
Google AI
on Orq.ai
Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Capabilities:
Chat
Code
Embeddings
Image Generation
Reasoning
Speech
Vision
Models Supported:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
Nano Banana 2 (gemini-3.1-flash-image)
Nano Banana 2 Lite (gemini-3.1-flash-lite-image)
Provider HQ:
Google (Alphabet Inc.), headquartered in Mountain View, California.
Access Google AI through Orq.ai’s AI Router
Google AI (via Gemini API / Google AI Studio) is Google’s managed platform for Gemini foundation models, offering text, code, image, and multimodal capabilities under a unified API. These models cover high‑end reasoning, general‑purpose coding, and cost‑efficient high‑volume use cases.
Orq.ai supports major Gemini variants (for example Gemini 3.1 Pro, Gemini 2.5 Flash, Gemini 2.0 Flash‑Lite), with availability depending on provider access, region, and your workspace configuration.
Google AI models available on Orq.ai
Model (example) | Type | Context | Best for | Pricing tier (reference) |
|---|---|---|---|---|
Gemini 3.1 Pro (or Gemini 2.5 Pro) | Chat / reasoning / vision | Long context (hundreds of thousands of tokens; check Google / Orq docs for current limits) | Complex reasoning, analysis, coding, and high‑stakes tasks where you want Google’s most capable Gemini model | Premium – Gemini 3.1 Pro pricing around 2.00 USD / 1M input tokens and 12.00 USD / 1M output tokens; Gemini 2.5 Pro around 1.25–2.50 USD input and 5–10+ USD output per 1M tokens depending on context tier |
Gemini 2.5 Flash | Chat / reasoning / coding | Large context (over 1M tokens in some tiers; verify in docs) | Everyday production workloads, RAG, coding, product features, and workflows that balance quality with cost and latency | Mid‑tier – Gemini 2.5 Flash typically around 0.30 USD / 1M input tokens and 2.50 USD / 1M output tokens, often cited as the best price–performance model in the Gemini family |
Gemini 2.0 Flash‑Lite / Gemini 2.5 Flash‑Lite | Chat / fast / cost‑efficient | Very large context (up to ~1M tokens; confirm in docs) | Fast, lower‑cost tasks such as classification, extraction, lightweight chat, routing, and high‑volume support workflows | Cost‑efficient – Gemini 2.0 Flash‑Lite and 2.5 Flash‑Lite tiers priced around 0.075–0.10 USD / 1M input tokens and 0.30–0.40 USD / 1M output tokens |
Pricing tiers here are approximate and based on Fal’s public model pricing; Orq.ai may apply its own billing or BYOK mapping. Always check Orq.ai’s pricing page and your configured Fal provider for current per‑model rates, quotas, and billing details.
Why use Google AI through Orq.ai?
Capability | Provider / models | Direct | Through Orq.ai |
|---|---|---|---|
Chat | Gemini chat models (for example Gemini 3.1 Pro, 2.5 Flash, 2.0 Flash‑Lite) | Call Gemini models directly via Google AI Studio / Gemini API for chat, reasoning, coding, and multimodal tasks | Use Gemini models through Orq.ai’s OpenAI‑compatible endpoint, adding routing, tracing, evals, budgets, and governance controls around each request |
Code | Gemini models tuned or suitable for coding/analysis | Use Gemini directly for code generation, debugging, refactoring, and agentic coding workflows | Route coding workloads through Orq.ai, compare Gemini models against other providers, and monitor cost, latency, and quality from one control layer. |
Embeddings / multimodal | Gemini models and dedicated embedding routes where available | Google AI provides multimodal support (text, image, audio, video) and some embeddings via Gemini | Use Orq.ai to route embedding or multimodal workloads to Gemini while keeping other providers for complementary tasks, all under unified observability |
This gives teams a practical way to use Google where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.
Pricing
Model rates
Gemini model pricing may differ depending on whether you:
connect your own Google AI / Gemini API key (BYOK), or
use Gemini models billed through Orq.ai where available
Key patterns:
High‑end models (Gemini 3.1 Pro, Gemini 2.5 Pro) sit in the 1.25–2.00 USD / 1M input and 5.00–12.00 USD / 1M output range depending on context tier.
Mid‑tier models (Gemini 2.5 Flash) sit around 0.30 USD / 1M input and 2.50 USD / 1M output.
Cheapest tiers (Gemini 2.0 Flash‑Lite, 2.5 Flash‑Lite) start as low as 0.075–0.10 USD / 1M input tokens and 0.30–0.40 USD / 1M output tokens.
Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details
Compatible frameworks and tools
Orq.ai exposes Google AI models through:
an OpenAI‑compatible API layer, and
native Google AI provider configuration in the AI Router.
That means:
Popular AI frameworks (for example, OpenAI‑compatible clients and orchestration libraries) can talk to Gemini via Orq’s router by pointing at Orq’s OpenAI‑compatible endpoint while Orq routes to Google AI behind the scenes.
Code assistants and IDE tools that support MCP or OpenAI‑compatible APIs (Cursor, VS Code, Claude Desktop, Warp, Zed, and similar tools you’ve documented) can route through Orq.ai to Gemini, depending on model and integration configuration.
Check the Orq.ai integration docs for the latest supported frameworks and tools for Google AI.
FAQs
Do I need a separate Google AI account to use Gemini through Orq.ai?
You can either connect your own Google AI / Gemini API key into Orq.ai or, where available, use Gemini models billed via Orq.ai; the exact options depend on your Orq plan, region, and how Google AI is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Google AI usage.
Can I route only some workflows to Google AI and others to different providers?
Yes. You define routes per workflow in Orq.ai and decide which ones should use Gemini vs other models, so you can reserve Google AI for specific regions, compliance needs, long‑context, or multimodal workloads while sending other tasks to different providers.
Does using Google AI through Orq.ai add latency?
Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets.
Alternatives to
Google AI
Anthropic
Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.
Chat
Code
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Image Generation
Reasoning
Speech
Vision
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
AWS
Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Reasoning
Vision
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5
Z.ai
Use Z.ai’s GLM‑5 family through a single Orq.ai integration. Route models such as GLM‑5.2, GLM‑5.1, GLM‑5, GLM‑5‑Turbo, GLM‑4.7, and GLM‑4.7‑FlashX via Orq’s AI Router for chat, reasoning, coding, multilingual tasks, vision, and cost‑efficient high‑volume workloads.
Chat
Image Generation
Reasoning
Vision
Models:
glm-5.2
glm-5.1
glm-5v-turbo


