
Cerebras
on Orq.ai
Use Cerebras‑hosted models through a single Orq.ai API. Route high‑performance open models such as GPT‑OSS, Llama, Qwen, and other supported models served by Cerebras Cloud through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Capabilities:
Chat
Reasoning
Vision
Models Supported:
cerebras/gemma-4-31b
zai-glm-4.7
cerebras/gpt-oss-120b
Provider HQ:
Cerebras Systems, headquartered in Sunnyvale, California.
Access Cerebras through Orq.ai’s AI Router
Cerebras provides an inference cloud optimized for fast, cost‑efficient serving of leading open and partner models, running on wafer‑scale systems with very high tokens‑per‑second throughput. These deployments cover high‑end reasoning, general‑purpose coding, and cost‑efficient high‑volume use cases
Orq.ai supports major Cerebras model offerings (for example GPT‑OSS‑class models and selected Llama/Qwen deployments), with availability depending on which models you enable under the Cerebras provider and your workspace configuration.
Cerebras models available on Orq.ai
Model | Type | Context / scope | Best for | Pricing tier |
|---|---|---|---|---|
GPT‑OSS‑class large model (for example gpt‑oss‑120B) | Chat / reasoning / vision | Large context (check Cerebras / Orq docs for current limits) | Complex reasoning, analysis, coding, and high‑stakes tasks where you want a large open model served at very high throughput on Cerebras Cloud | Premium – Cerebras Inference higher tiers; typically positioned above smaller models, with pay‑per‑token or pay‑per‑model options |
Mid‑sized open model (for example Llama 3.x 8B–70B, Qwen 2.5 / 3 mid‑tier) | Chat / reasoning / coding | Medium–large context (verify in Cerebras / Orq docs) | Everyday production workloads, RAG, coding, product features, and workflows that balance quality with speed and cost | Mid‑tier – Cerebras Developer / paid plans; often in the ~0.10–0.60 USD per million token range depending on model choice |
Smaller / fast open model (for example Llama 3.1 8B, lighter GPT‑OSS tiers) | Chat / fast / cost‑efficient | Medium context (confirm in docs) | Fast, lower‑cost tasks such as classification, extraction, lightweight chat, routing, and high‑volume support workflows | Cost‑efficient – Cerebras Developer tier pricing; lower per‑million token cost designed for high‑throughput workloads and experimentation |
Pricing tiers here are approximate and based on Cerebras’s public Inference Cloud pricing; Orq.ai may apply its own billing or BYOK mapping. Always check Orq.ai’s pricing page and your configured Cerebras provider for current per‑model rates, quotas, and billing details.
Why use Cerebras through Orq.ai?
Capability | Provider / models | Direct | Through Orq.ai |
|---|---|---|---|
Chat | Cerebras‑served chat models (for example GPT‑OSS, Llama, Qwen on Cerebras) | Call Cerebras models directly via the Cerebras Inference API for chat, reasoning, coding, and multimodal tasks | Use Cerebras models through Orq.ai’s OpenAI‑compatible endpoint, adding routing, tracing, evals, budgets, and governance controls around each request |
Code | Open models on Cerebras tuned or suitable for coding/analysis | Use Cerebras directly for code generation, debugging, refactoring, and agentic coding workflows, often with higher tokens‑per‑second vs GPU clouds | Route coding workloads through Orq.ai, compare Cerebras‑served models against other providers, and monitor cost, latency, and quality from one control layer |
Embeddings | Embedding providers configured in Orq.ai | Cerebras focuses on high‑performance serving of LLMs; embeddings may be handled by separate models or providers | Use Orq.ai to route embedding workloads to supported embedding providers while keeping Cerebras for reasoning, generation, and agent steps |
This gives teams a practical way to use Cerebras where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.
Pricing
Tier / model | Input / price | Output / included | Cached / notes |
|---|---|---|---|
Model rates | Cerebras model pricing may differ depending on whether you: | * connect your own Cerebras Cloud account and API key (BYOK), or | * use Cerebras‑hosted models billed through Orq.ai where available Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details. |
Compatible frameworks and tools
Orq.ai exposes Cerebras models through:
an OpenAI‑compatible API layer, and
native SDKs and router integrations where applicable
That means:
Popular AI frameworks (for example, OpenAI‑compatible clients and orchestration libraries) can talk to Cerebras models via Orq’s router.
Code assistants and IDE tools that support MCP or OpenAI‑compatible APIs (Cursor, VS Code, Claude Desktop, Warp, Zed, and similar tools you’ve already documented) can route through Orq.ai to Cerebras, depending on model and integration configuration.
Check the Orq.ai integration docs for the latest supported frameworks and tools for Cerebras.
FAQs
Do I need a separate Cerebras account to use Cerebras through Orq.ai?
You can either connect your own Cerebras API key into Orq.ai (for example from the Cerebras Developer Console or Inference Cloud) or, where available, use Cerebras‑hosted models billed via Orq.ai; the exact options depend on your Orq plan, region, and how Cerebras is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Cerebras usage.
Can I route only some workflows to Cerebras and others to different providers?
Yes. You define routes per workflow in Orq.ai and decide which ones should use Cerebras vs other providers, so you can reserve Cerebras for specific performance‑sensitive or open‑model workloads while sending other tasks to different models.
Does using Cerebras through Orq.ai add latency?
Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets, while still benefiting from Cerebras’s high tokens‑per‑second capabilities.
Alternatives to
Cerebras
Anthropic
Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.
Chat
Code
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Image Generation
Reasoning
Speech
Vision
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
Google AI
Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Image Generation
Reasoning
Speech
Vision
Models:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
AWS
Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Reasoning
Vision
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5


