Contextual AI

on Orq.ai

Use Contextual AI’s models through a single Orq.ai API. Route models such as Generate, Rerank, and LMUnit via Orq’s AI Router for grounded chat, reasoning, evaluation, and retrieval‑augmented workloads.

Capabilities:

Models Supported:

No models available

Provider HQ:

Contextual AI, headquartered in Mountain View, California, United States.

Access Contextual AI through Orq.ai’s AI Router

Contextual AI builds foundation models and tools optimized for trustworthy, grounded enterprise AI, including a Generate LLM focused on minimizing hallucinations, a Rerank model for retrieval, and LMUnit for evaluation and scoring. These models cover high‑end grounded generation, ranking, and evaluation for RAG and enterprise applications.

Orq.ai supports Contextual AI’s main model types (Generate, Rerank, LMUnit), with availability depending on provider access, region, and your workspace configuration.

Contextual AI models available on Orq.ai

Model

Type

Context / scope

Best for

Pricing tier (reference)

Generate

Chat / grounded generation

LLM context window (check Contextual / Orq docs for current limits)

Grounded, low‑hallucination generation for enterprise RAG, chat, and workflow automation

Premium – priced around 3 USD / 1M input tokens and 15 USD / 1M output tokens for Generate, reflecting a high‑quality grounded model

Rerank (v2)

Reranking / retrieval control

Operates over retrieved documents / candidates

Improving search and RAG by reordering retrieved items based on instruction‑following priorities

Mid‑tier – Rerank‑v2 currently around 0.05 USD per 1M tokens; Rerank‑v2‑mini around 0.02 USD per 1M tokens, making it cost‑efficient for high‑volume retrieval

LMUnit

Evaluation / scoring

Evaluation tasks and test suites

Preference modeling, direct scoring, natural‑language unit test evaluation, and LLM‑powered eval pipelines

Evaluation tier – LMUnit priced around 3 USD / 1M input tokens; typically used selectively for eval and QA rather than every request

Pricing tiers here are based on Contextual AI’s public pricing; Orq.ai may apply its own billing or BYOK mapping. Always check Orq.ai’s pricing page and your configured Contextual AI provider for current per‑model rates, quotas, and billing details.

Why use Contextual AI through Orq.ai?

Capability

Provider / models

Direct

Through Orq.ai

Chat / RAG

Generate (grounded LLM)

Call Generate directly via Contextual AI’s API for grounded chat and RAG‑style generation that minimizes hallucinations

Use Generate through Orq.ai’s OpenAI‑compatible endpoint, adding routing, tracing, evals, budgets, and governance controls around each request

Retrieval

Rerank (v2, v2‑mini)

Use Rerank directly to reorder search/RAG results based on instructions and relevance

Route retrieval scoring through Orq.ai while keeping unified logs, metrics, and experiments across multiple rerank or search providers

Evaluation

LMUnit (evaluation‑optimized model)

Use LMUnit directly for preference tests, scoring, and unit‑test‑style evaluations of prompts and outputs

Integrate LMUnit into Orq.ai Skills and eval pipelines, running centralized evaluations across providers and workflows from one control plane

This gives teams a practical way to use Contextual AI where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Model rates

Contextual AI model pricing may differ depending on whether you:

  • connect your own Contextual AI account and API key (BYOK), or

  • use Contextual AI models billed through Orq.ai where available

Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details.

Compatible frameworks and tools

Orq.ai exposes Contextual AI models through:

  • an OpenAI‑compatible API layer for generation, and

  • native or HTTP integrations for Rerank and evaluation workflows where applicable

That means:

Popular AI frameworks (for example, OpenAI‑compatible clients and orchestration libraries) can talk to Generate via Orq’s router.

Retrieval and evaluation pipelines can integrate Rerank and LMUnit via Orq Skills and eval runners, keeping everything under a single observability and governance layer.

Check the Orq.ai integration docs for the latest supported frameworks and tools for Contextual AI.

FAQs

Do I need a separate Contextual AI account to use Contextual AI through Orq.ai?

You can either connect your own Contextual AI API key into Orq.ai or, where available, use Contextual AI models billed via Orq.ai; the exact options depend on your Orq plan, region, and how Contextual AI is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Contextual AI usage.

Can I route only some workflows to Contextual AI and others to different providers?

Yes. You define routes per workflow in Orq.ai and decide which ones should use Generate/Rerank/LMUnit vs other providers, so you can reserve Contextual AI for grounded RAG and evaluation paths while sending simpler or different workloads elsewhere.

Does using Contextual AI through Orq.ai add latency?

Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets.

Alternatives to

Contextual AI

Anthropic

Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.

Chat

Code

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI

Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Image Generation

Reasoning

Speech

Vision

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI

Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Image Generation

Reasoning

Speech

Vision

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS

Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Reasoning

Vision

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.