
OpenAI-compatible
on Orq.ai
Use OpenAI‑compatible models through a single Orq.ai API. Route traffic to providers like Groq, Together AI, Mistral, Moonshot, and any custom OpenAI‑compatible deployment via Orq’s AI Router, while keeping your existing OpenAI SDK and request shapes.
Capabilities:
Models Supported:
No models available
Provider HQ:
Varies per underlying provider; “OpenAI‑compatible” here refers to the API format, not a single vendor.
Access OpenAI-compatible through Orq.ai’s AI Router
OpenAI‑compatible endpoints are APIs that follow the OpenAI format (paths, payloads, and responses) but are backed by different providers or your own hosted models. Orq.ai’s AI Gateway exposes an OpenAI‑compatible base URL so applications can talk to hundreds of models and providers with the same client libraries and JSON schemas.
OpenAI-compatible models available on Orq.ai
Model | Type | Context / scope | Best for | Pricing tier |
|---|---|---|---|---|
Groq Llama‑3.3‑70B via OpenAI‑compatible | Chat / reasoning / coding | Around 128K context (check Groq / Orq docs) | Complex reasoning and coding on a large open‑weight model served at very low latency through a Groq OpenAI‑style endpoint | Premium – public data shows roughly 0.59 USD / 1M input tokens and 0.79 USD / 1M output tokens, with discounts for prompt caching and batch |
Mistral Large 3 via OpenAI‑compatible | Chat / reasoning / coding | Long context (hundreds of thousands of tokens; verify per provider) | Everyday production workloads and RAG where you want Mistral’s flagship model under the OpenAI format for easier integration | Mid‑tier – many sources list around 0.50 USD / 1M input tokens and 1.50 USD / 1M output tokens, depending on the upstream Mistral provider you configure |
Moonshot / Together / other OpenAI‑like models | Chat / fast / cost‑efficient | Medium–long context (for example 32K–256K tokens depending on model) | Fast, lower‑cost tasks such as classification, extraction, lightweight chat, routing, and high‑volume workloads across different regions and vendors | Cost‑efficient – many OpenAI‑compatible endpoints start in the 0.02–0.10 USD / 1M input tokens range for small or distilled models, rising into mid‑tier prices for larger ones |
Why use OpenAI-compatible through Orq.ai?
Capability | Provider / models | Direct | Through Orq.ai |
|---|---|---|---|
Chat | Any OpenAI‑compatible chat/completions endpoint (Groq, Mistral, Moonshot, Together, custom) | Call each provider’s OpenAI‑style endpoint separately, manage multiple base URLs and keys, and handle routing/fallbacks yourself | Use Orq.ai’s single OpenAI‑compatible base URL (https://api.orq.ai/v3/router) to route to any configured provider, with centralised routing, tracing, evals, budgets, and governance controls |
Code | OpenAI‑compatible code‑friendly models (for example Llama 3.x, Mistral Small, Kimi K2.x via routers) | Use providers directly for code generation, debugging, and refactoring via their OpenAI‑format APIs | Route coding workloads through Orq.ai, compare multiple OpenAI‑compatible backends behind one endpoint, and monitor cost, latency, and quality from a single control layer |
Vision / multimodal | OpenAI‑compatible endpoints that support images, audio, and embeddings | Call each multimodal or embedding endpoint directly, wiring each one into your app manually | Use Orq.ai’s OpenAI‑compatible router so chat, vision, audio, and embeddings all share the same base URL, benefiting from fallbacks, caching, load‑balancing, and full observability |
This gives teams a practical way to use any OpenAI-compatible provider where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.
Pricing
Plans and API access
OpenAI‑compatible usage in Orq.ai depends on:
which providers you connect (for example Groq, Mistral, Moonshot, Together, custom), and
whether you use BYOK (your own provider keys) or usage billed through Orq.ai where available
In general:
Entry / cheapest OpenAI‑compatible models: small/distilled models often start around 0.02–0.10 USD / 1M input tokens
Mid‑tier: mainstream 7–32B models commonly land in the 0.10–0.60 USD / 1M input tokens range
Premium: large or long‑context models (for example Llama‑3.3‑70B, Mistral Large) can run from 0.50–2.00 USD / 1M input tokens, with higher output rates
Some providers also offer Batch/Flex or caching discounts, which Orq.ai can help you take advantage of by centralising routing policies.
Check the Orq.ai pricing page and your workspace’s provider configuration for current options around using OpenAI‑compatible endpoints via Orq.ai.
Compatible frameworks and tools
Orq.ai exposes OpenAI‑compatible routing through:
the OpenAI‑compatible base URL https://api.orq.ai/v3/router in the AI Gateway, and
OpenAI‑like provider entries in Model Garden / AI Router for upstreams such as Groq, Together, Mistral, Moonshot, and your own deployments
That means:
Existing backend workflows using OpenAI SDKs or other OpenAI‑compatible clients can switch to Orq.ai by changing only the base URL and the API key, without altering request payloads or client code
Agents, code assistants, and tools that already integrate with Orq.ai (for example via OpenAI‑compatible or HTTP tools) can route through Orq’s endpoint and still reach non‑OpenAI providers under the hood
Check the Orq.ai integration docs for the latest supported frameworks and tools for OpenAI‑compatible models.
FAQs
Do I need separate accounts for each OpenAI‑compatible provider to use them through Orq.ai?
Yes. You connect each upstream provider (Groq, Mistral, Moonshot, Together, etc.) with its own API key into Orq.ai, or in some cases use usage billed via Orq.ai; the exact options depend on your Orq plan, region, and how each provider is configured. In all cases, Orq.ai gives you one place to manage routing, observability, and cost controls around your OpenAI‑compatible usage.
Can I route only some workflows to OpenAI‑compatible providers and others to native APIs?
Yes. You define routes per workflow in Orq.ai and decide which ones should use OpenAI‑compatible endpoints vs native provider integrations, so you can experiment with multiple backends without changing application code.
Does using OpenAI‑compatible through Orq.ai add latency?
Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets while gaining visibility and control.
Alternatives to
OpenAI-compatible
Anthropic
Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.
Chat
Code
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Image Generation
Reasoning
Speech
Vision
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
Google AI
Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Image Generation
Reasoning
Speech
Vision
Models:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
AWS
Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Reasoning
Vision
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5


