OpenAI-compatible

on Orq.ai

Use OpenAI‑compatible models through a single Orq.ai API. Route traffic to providers like Groq, Together AI, Mistral, Moonshot, and any custom OpenAI‑compatible deployment via Orq’s AI Router, while keeping your existing OpenAI SDK and request shapes.

Capabilities:

Models Supported:

No models available

Provider HQ:

Varies per underlying provider; “OpenAI‑compatible” here refers to the API format, not a single vendor.

Access OpenAI-compatible through Orq.ai’s AI Router

OpenAI‑compatible endpoints are APIs that follow the OpenAI format (paths, payloads, and responses) but are backed by different providers or your own hosted models. Orq.ai’s AI Gateway exposes an OpenAI‑compatible base URL so applications can talk to hundreds of models and providers with the same client libraries and JSON schemas.

OpenAI-compatible models available on Orq.ai

Model

Type

Context / scope

Best for

Pricing tier

Groq Llama‑3.3‑70B via OpenAI‑compatible

Chat / reasoning / coding

Around 128K context (check Groq / Orq docs)

Complex reasoning and coding on a large open‑weight model served at very low latency through a Groq OpenAI‑style endpoint

Premium – public data shows roughly 0.59 USD / 1M input tokens and 0.79 USD / 1M output tokens, with discounts for prompt caching and batch

Mistral Large 3 via OpenAI‑compatible

Chat / reasoning / coding

Long context (hundreds of thousands of tokens; verify per provider)

Everyday production workloads and RAG where you want Mistral’s flagship model under the OpenAI format for easier integration

Mid‑tier – many sources list around 0.50 USD / 1M input tokens and 1.50 USD / 1M output tokens, depending on the upstream Mistral provider you configure

Moonshot / Together / other OpenAI‑like models

Chat / fast / cost‑efficient

Medium–long context (for example 32K–256K tokens depending on model)

Fast, lower‑cost tasks such as classification, extraction, lightweight chat, routing, and high‑volume workloads across different regions and vendors

Cost‑efficient – many OpenAI‑compatible endpoints start in the 0.02–0.10 USD / 1M input tokens range for small or distilled models, rising into mid‑tier prices for larger ones

Why use OpenAI-compatible through Orq.ai?

Capability

Provider / models

Direct

Through Orq.ai

Chat

Any OpenAI‑compatible chat/completions endpoint (Groq, Mistral, Moonshot, Together, custom)

Call each provider’s OpenAI‑style endpoint separately, manage multiple base URLs and keys, and handle routing/fallbacks yourself

Use Orq.ai’s single OpenAI‑compatible base URL (https://api.orq.ai/v3/router) to route to any configured provider, with centralised routing, tracing, evals, budgets, and governance controls

Code

OpenAI‑compatible code‑friendly models (for example Llama 3.x, Mistral Small, Kimi K2.x via routers)

Use providers directly for code generation, debugging, and refactoring via their OpenAI‑format APIs

Route coding workloads through Orq.ai, compare multiple OpenAI‑compatible backends behind one endpoint, and monitor cost, latency, and quality from a single control layer

Vision / multimodal

OpenAI‑compatible endpoints that support images, audio, and embeddings

Call each multimodal or embedding endpoint directly, wiring each one into your app manually

Use Orq.ai’s OpenAI‑compatible router so chat, vision, audio, and embeddings all share the same base URL, benefiting from fallbacks, caching, load‑balancing, and full observability

This gives teams a practical way to use any OpenAI-compatible provider where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Plans and API access

OpenAI‑compatible usage in Orq.ai depends on:

  • which providers you connect (for example Groq, Mistral, Moonshot, Together, custom), and

  • whether you use BYOK (your own provider keys) or usage billed through Orq.ai where available

In general:

  • Entry / cheapest OpenAI‑compatible models: small/distilled models often start around 0.02–0.10 USD / 1M input tokens

  • Mid‑tier: mainstream 7–32B models commonly land in the 0.10–0.60 USD / 1M input tokens range

  • Premium: large or long‑context models (for example Llama‑3.3‑70B, Mistral Large) can run from 0.50–2.00 USD / 1M input tokens, with higher output rates

Some providers also offer Batch/Flex or caching discounts, which Orq.ai can help you take advantage of by centralising routing policies.

Check the Orq.ai pricing page and your workspace’s provider configuration for current options around using OpenAI‑compatible endpoints via Orq.ai.

Compatible frameworks and tools

Orq.ai exposes OpenAI‑compatible routing through:

  • the OpenAI‑compatible base URL https://api.orq.ai/v3/router in the AI Gateway, and

  • OpenAI‑like provider entries in Model Garden / AI Router for upstreams such as Groq, Together, Mistral, Moonshot, and your own deployments

That means:

  • Existing backend workflows using OpenAI SDKs or other OpenAI‑compatible clients can switch to Orq.ai by changing only the base URL and the API key, without altering request payloads or client code

  • Agents, code assistants, and tools that already integrate with Orq.ai (for example via OpenAI‑compatible or HTTP tools) can route through Orq’s endpoint and still reach non‑OpenAI providers under the hood

Check the Orq.ai integration docs for the latest supported frameworks and tools for OpenAI‑compatible models.

FAQs

Do I need separate accounts for each OpenAI‑compatible provider to use them through Orq.ai?

Yes. You connect each upstream provider (Groq, Mistral, Moonshot, Together, etc.) with its own API key into Orq.ai, or in some cases use usage billed via Orq.ai; the exact options depend on your Orq plan, region, and how each provider is configured. In all cases, Orq.ai gives you one place to manage routing, observability, and cost controls around your OpenAI‑compatible usage.

Can I route only some workflows to OpenAI‑compatible providers and others to native APIs?

Yes. You define routes per workflow in Orq.ai and decide which ones should use OpenAI‑compatible endpoints vs native provider integrations, so you can experiment with multiple backends without changing application code.

Does using OpenAI‑compatible through Orq.ai add latency?

Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets while gaining visibility and control.

Alternatives to

OpenAI-compatible

Anthropic

Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.

Chat

Code

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI

Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Image Generation

Reasoning

Speech

Vision

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI

Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Image Generation

Reasoning

Speech

Vision

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS

Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Reasoning

Vision

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.