Google AI

on Orq.ai

Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Capabilities:

Chat

Code

Embeddings

Image Generation

Reasoning

Speech

Vision

Models Supported:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

Nano Banana 2 (gemini-3.1-flash-image)

Nano Banana 2 Lite (gemini-3.1-flash-lite-image)

Provider HQ:

Google (Alphabet Inc.), headquartered in Mountain View, California.

Access Google AI through Orq.ai’s AI Router

Google AI (via Gemini API / Google AI Studio) is Google’s managed platform for Gemini foundation models, offering text, code, image, and multimodal capabilities under a unified API. These models cover high‑end reasoning, general‑purpose coding, and cost‑efficient high‑volume use cases.

Orq.ai supports major Gemini variants (for example Gemini 3.1 Pro, Gemini 2.5 Flash, Gemini 2.0 Flash‑Lite), with availability depending on provider access, region, and your workspace configuration.

Google AI models available on Orq.ai

Model (example)

Type

Context

Best for

Pricing tier (reference)

Gemini 3.1 Pro (or Gemini 2.5 Pro)

Chat / reasoning / vision

Long context (hundreds of thousands of tokens; check Google / Orq docs for current limits)

Complex reasoning, analysis, coding, and high‑stakes tasks where you want Google’s most capable Gemini model

Premium – Gemini 3.1 Pro pricing around 2.00 USD / 1M input tokens and 12.00 USD / 1M output tokens; Gemini 2.5 Pro around 1.25–2.50 USD input and 5–10+ USD output per 1M tokens depending on context tier

Gemini 2.5 Flash

Chat / reasoning / coding

Large context (over 1M tokens in some tiers; verify in docs)

Everyday production workloads, RAG, coding, product features, and workflows that balance quality with cost and latency

Mid‑tier – Gemini 2.5 Flash typically around 0.30 USD / 1M input tokens and 2.50 USD / 1M output tokens, often cited as the best price–performance model in the Gemini family

Gemini 2.0 Flash‑Lite / Gemini 2.5 Flash‑Lite

Chat / fast / cost‑efficient

Very large context (up to ~1M tokens; confirm in docs)

Fast, lower‑cost tasks such as classification, extraction, lightweight chat, routing, and high‑volume support workflows

Cost‑efficient – Gemini 2.0 Flash‑Lite and 2.5 Flash‑Lite tiers priced around 0.075–0.10 USD / 1M input tokens and 0.30–0.40 USD / 1M output tokens

Pricing tiers here are approximate and based on Fal’s public model pricing; Orq.ai may apply its own billing or BYOK mapping. Always check Orq.ai’s pricing page and your configured Fal provider for current per‑model rates, quotas, and billing details.

Why use Google AI through Orq.ai?

Capability

Provider / models

Direct

Through Orq.ai

Chat

Gemini chat models (for example Gemini 3.1 Pro, 2.5 Flash, 2.0 Flash‑Lite)

Call Gemini models directly via Google AI Studio / Gemini API for chat, reasoning, coding, and multimodal tasks

Use Gemini models through Orq.ai’s OpenAI‑compatible endpoint, adding routing, tracing, evals, budgets, and governance controls around each request

Code

Gemini models tuned or suitable for coding/analysis

Use Gemini directly for code generation, debugging, refactoring, and agentic coding workflows

Route coding workloads through Orq.ai, compare Gemini models against other providers, and monitor cost, latency, and quality from one control layer.

Embeddings / multimodal

Gemini models and dedicated embedding routes where available

Google AI provides multimodal support (text, image, audio, video) and some embeddings via Gemini

Use Orq.ai to route embedding or multimodal workloads to Gemini while keeping other providers for complementary tasks, all under unified observability

This gives teams a practical way to use Google where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Model rates

Gemini model pricing may differ depending on whether you:

  • connect your own Google AI / Gemini API key (BYOK), or

  • use Gemini models billed through Orq.ai where available

Key patterns:

  • High‑end models (Gemini 3.1 Pro, Gemini 2.5 Pro) sit in the 1.25–2.00 USD / 1M input and 5.00–12.00 USD / 1M output range depending on context tier.

  • Mid‑tier models (Gemini 2.5 Flash) sit around 0.30 USD / 1M input and 2.50 USD / 1M output.

  • Cheapest tiers (Gemini 2.0 Flash‑Lite, 2.5 Flash‑Lite) start as low as 0.075–0.10 USD / 1M input tokens and 0.30–0.40 USD / 1M output tokens.

Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details

Compatible frameworks and tools

Orq.ai exposes Google AI models through:

  • an OpenAI‑compatible API layer, and

  • native Google AI provider configuration in the AI Router.

That means:

  • Popular AI frameworks (for example, OpenAI‑compatible clients and orchestration libraries) can talk to Gemini via Orq’s router by pointing at Orq’s OpenAI‑compatible endpoint while Orq routes to Google AI behind the scenes.

  • Code assistants and IDE tools that support MCP or OpenAI‑compatible APIs (Cursor, VS Code, Claude Desktop, Warp, Zed, and similar tools you’ve documented) can route through Orq.ai to Gemini, depending on model and integration configuration.

Check the Orq.ai integration docs for the latest supported frameworks and tools for Google AI.

FAQs

Do I need a separate Google AI account to use Gemini through Orq.ai?

You can either connect your own Google AI / Gemini API key into Orq.ai or, where available, use Gemini models billed via Orq.ai; the exact options depend on your Orq plan, region, and how Google AI is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Google AI usage.

Can I route only some workflows to Google AI and others to different providers?

Yes. You define routes per workflow in Orq.ai and decide which ones should use Gemini vs other models, so you can reserve Google AI for specific regions, compliance needs, long‑context, or multimodal workloads while sending other tasks to different providers.

Does using Google AI through Orq.ai add latency?

Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets.

Alternatives to

Google AI

Anthropic

Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.

Chat

Code

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI

Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Image Generation

Reasoning

Speech

Vision

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

AWS

Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Reasoning

Vision

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Z.ai

Use Z.ai’s GLM‑5 family through a single Orq.ai integration. Route models such as GLM‑5.2, GLM‑5.1, GLM‑5, GLM‑5‑Turbo, GLM‑4.7, and GLM‑4.7‑FlashX via Orq’s AI Router for chat, reasoning, coding, multilingual tasks, vision, and cost‑efficient high‑volume workloads.

Chat

Image Generation

Reasoning

Vision

Models:

glm-5.2

glm-5.1

glm-5v-turbo

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.