Vertex AI

on Orq.ai

Use Google Vertex AI’s Gemini models through a single Orq.ai API. Route models such as Gemini 2.5 Pro, Gemini 2.5 Flash, and Gemini 1.5 Flash via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Capabilities:

Models Supported:

No models available

Provider HQ:

Google (Google Cloud / Gemini Enterprise Agent Platform), headquartered in Mountain View, California.

Access Google Vertex AI through Orq.ai’s AI Router

Vertex AI (now part of the Gemini Enterprise Agent Platform) is Google Cloud’s managed platform for foundation models, offering Gemini text and multimodal models plus additional third‑party and open models through Model Garden. These models cover high‑end reasoning, general‑purpose coding, and cost‑efficient high‑volume use cases.

Orq.ai supports major Gemini variants on Vertex AI (for example Gemini 2.5 Pro, Gemini 2.5 Flash, Gemini 1.5 Flash), with availability depending on project setup, region, and your workspace configuration.

Google Vertex AI models available on Orq.ai

Model (example)

Type

Context

Best for

Pricing tier (reference)

Gemini 2.5 Pro (Vertex AI)

Chat / reasoning / vision

Up to around 1–2M tokens (check Vertex / Orq docs for current limits)

Complex reasoning, analysis, coding, and high‑stakes tasks where you want Google Cloud’s most capable Gemini model in your region

Premium – Gemini 2.5 Pro on Vertex AI priced around 1.25 USD / 1M input tokens and 10.00 USD / 1M output tokens for base context, rising to about 2.50 input / 15.00 output at higher context tiers

Gemini 1.5 Pro / Gemini 3.1 Pro (where available)

Chat / reasoning / coding

Long context (up to 2M tokens)

Large‑context RAG, long documents, multi‑step workflows, and enterprise applications needing long‑context reasoning

Alternative premium – Gemini 1.5 Pro priced roughly 1.25–2.50 USD / 1M input and 5–10 USD / 1M output depending on context tier

Gemini 2.5 Flash (Vertex AI)

Chat / reasoning / coding

Around 1M tokens

Everyday production workloads, RAG, coding, product features, and workflows that balance quality with cost and latency

Mid‑tier – Gemini 2.5 Flash around 0.30 USD / 1M input tokens and 2.50 USD / 1M output tokens

Gemini 1.5 Flash / Flash‑8B

Chat / fast / cost‑efficient

About 1M tokens

Fast, lower‑cost tasks such as classification, extraction, lightweight chat, routing, and high‑volume support workflows

Cost‑efficient – Gemini 1.5 Flash around 0.075 USD / 1M input and 0.30 USD / 1M output tokens; Gemini 1.5 Flash‑8B around 0.037 USD input and 0.15 USD output per 1M tokens

Pricing tiers here are approximate and based on Vertex AI’s public Gemini pricing; Orq.ai may apply its own billing or BYOK mapping. Always check Orq.ai’s pricing page and your configured Vertex provider for current per‑model rates, quotas, and billing details.

Why use Google Vertex AI through Orq.ai?

Capability

Provider / models

Direct

Through Orq.ai

Chat

Gemini chat models on Vertex (for example Gemini 2.5 Pro, 2.5 Flash, 1.5 Flash)

Call Gemini models directly via Vertex AI’s generative AI APIs for chat, reasoning, coding, and multimodal tasks

Use Vertex‑hosted Gemini models through Orq.ai’s OpenAI‑compatible endpoint, adding routing, tracing, evals, budgets, and governance controls around each request

Code

Gemini models tuned or suitable for coding/analysis

Use Gemini on Vertex AI for code generation, debugging, refactoring, and agentic coding workflows

Route coding workloads through Orq.ai, compare Gemini on Vertex against other providers, and monitor cost, latency, and quality from one control layer

Embeddings / multimodal

Gemini models and Vertex embedding endpoints

Vertex AI offers Gemini text, image, and multimodal support, as well as embedding models such as Embedding 003

Use Orq.ai to route embedding or multimodal workloads to Vertex while keeping other providers for complementary tasks, with unified observability

This gives teams a practical way to use Vertex-hosted models where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Model rates

Gemini on Vertex AI pricing may differ depending on whether you:

  • connect your own Google Cloud project and Vertex AI (BYOK), or

  • use Vertex models billed through Orq.ai where available

Key patterns from current public pricing:

  • Gemini 2.5 Pro: about 1.25 USD / 1M input and 10.00 USD / 1M output tokens (higher at larger context tiers).

  • Gemini 2.5 Flash: about 0.30 USD / 1M input and 2.50 USD / 1M output tokens.

  • Gemini 1.5 Flash: about 0.075 USD / 1M input and 0.30 USD / 1M output tokens.

  • Embedding 003: roughly 0.02 USD per 1M tokens for both input and output

Vertex AI also charges separately for training, custom deployment, and certain grounding features, but those costs typically sit outside pure token usage.

Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details.

Compatible frameworks and tools

Orq.ai exposes Vertex AI through:

  • a dedicated Vertex AI provider configuration in the Model Garden / AI Router, and

  • an OpenAI‑compatible API layer for applications that expect that interface.

That means:

Popular AI frameworks (for example, OpenAI‑compatible clients and orchestration libraries) can talk to Gemini on Vertex via Orq’s router by pointing at Orq’s OpenAI‑compatible endpoint while Orq routes to Vertex behind the scenes.

Code assistants and IDE tools that support MCP or OpenAI‑compatible APIs (Cursor, VS Code, Claude Desktop, Warp, Zed, and similar tools you’ve documented) can route through Orq.ai to Gemini on Vertex, depending on model and integration configuration.

Check the Orq.ai integration docs for the latest supported frameworks and tools for Vertex AI.

FAQs

Do I need a separate Google Cloud account to use Vertex AI through Orq.ai?

You can either connect your own Google Cloud project, Vertex AI configuration, and service account into Orq.ai or, where available, use Vertex models billed via Orq.ai; the exact options depend on your Orq plan, region, and how Vertex is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Vertex AI usage.

Can I route only some workflows to Vertex AI and others to different providers?

Yes. You define routes per workflow in Orq.ai and decide which ones should use Gemini on Vertex vs other models, so you can reserve Vertex for specific regions, compliance needs, or long‑context / multimodal workloads while sending other tasks to different providers.

Does using Vertex AI through Orq.ai add latency?

Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets.

Alternatives to

Vertex AI

Anthropic

Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.

Chat

Code

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI

Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Image Generation

Reasoning

Speech

Vision

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI

Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Image Generation

Reasoning

Speech

Vision

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS

Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Reasoning

Vision

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.