Cerebras

on Orq.ai

Use Cerebras‑hosted models through a single Orq.ai API. Route high‑performance open models such as GPT‑OSS, Llama, Qwen, and other supported models served by Cerebras Cloud through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Capabilities:

Chat

Reasoning

Vision

Models Supported:

cerebras/gemma-4-31b

zai-glm-4.7

cerebras/gpt-oss-120b

Provider HQ:

Cerebras Systems, headquartered in Sunnyvale, California.

Access Cerebras through Orq.ai’s AI Router

Cerebras provides an inference cloud optimized for fast, cost‑efficient serving of leading open and partner models, running on wafer‑scale systems with very high tokens‑per‑second throughput. These deployments cover high‑end reasoning, general‑purpose coding, and cost‑efficient high‑volume use cases

Orq.ai supports major Cerebras model offerings (for example GPT‑OSS‑class models and selected Llama/Qwen deployments), with availability depending on which models you enable under the Cerebras provider and your workspace configuration.

Cerebras models available on Orq.ai

Model

Type

Context / scope

Best for

Pricing tier

GPT‑OSS‑class large model (for example gpt‑oss‑120B)

Chat / reasoning / vision

Large context (check Cerebras / Orq docs for current limits)

Complex reasoning, analysis, coding, and high‑stakes tasks where you want a large open model served at very high throughput on Cerebras Cloud

Premium – Cerebras Inference higher tiers; typically positioned above smaller models, with pay‑per‑token or pay‑per‑model options

Mid‑sized open model (for example Llama 3.x 8B–70B, Qwen 2.5 / 3 mid‑tier)

Chat / reasoning / coding

Medium–large context (verify in Cerebras / Orq docs)

Everyday production workloads, RAG, coding, product features, and workflows that balance quality with speed and cost

Mid‑tier – Cerebras Developer / paid plans; often in the ~0.10–0.60 USD per million token range depending on model choice

Smaller / fast open model (for example Llama 3.1 8B, lighter GPT‑OSS tiers)

Chat / fast / cost‑efficient

Medium context (confirm in docs)

Fast, lower‑cost tasks such as classification, extraction, lightweight chat, routing, and high‑volume support workflows

Cost‑efficient – Cerebras Developer tier pricing; lower per‑million token cost designed for high‑throughput workloads and experimentation

Pricing tiers here are approximate and based on Cerebras’s public Inference Cloud pricing; Orq.ai may apply its own billing or BYOK mapping. Always check Orq.ai’s pricing page and your configured Cerebras provider for current per‑model rates, quotas, and billing details.

Why use Cerebras through Orq.ai?

Capability

Provider / models

Direct

Through Orq.ai

Chat

Cerebras‑served chat models (for example GPT‑OSS, Llama, Qwen on Cerebras)

Call Cerebras models directly via the Cerebras Inference API for chat, reasoning, coding, and multimodal tasks

Use Cerebras models through Orq.ai’s OpenAI‑compatible endpoint, adding routing, tracing, evals, budgets, and governance controls around each request

Code

Open models on Cerebras tuned or suitable for coding/analysis

Use Cerebras directly for code generation, debugging, refactoring, and agentic coding workflows, often with higher tokens‑per‑second vs GPU clouds

Route coding workloads through Orq.ai, compare Cerebras‑served models against other providers, and monitor cost, latency, and quality from one control layer

Embeddings

Embedding providers configured in Orq.ai

Cerebras focuses on high‑performance serving of LLMs; embeddings may be handled by separate models or providers

Use Orq.ai to route embedding workloads to supported embedding providers while keeping Cerebras for reasoning, generation, and agent steps

This gives teams a practical way to use Cerebras where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Tier / model

Input / price

Output / included

Cached / notes

Model rates

Cerebras model pricing may differ depending on whether you:

* connect your own Cerebras Cloud account and API key (BYOK), or

* use Cerebras‑hosted models billed through Orq.ai where available Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details.

Compatible frameworks and tools

Orq.ai exposes Cerebras models through:

  • an OpenAI‑compatible API layer, and

  • native SDKs and router integrations where applicable

That means:

  • Popular AI frameworks (for example, OpenAI‑compatible clients and orchestration libraries) can talk to Cerebras models via Orq’s router.

  • Code assistants and IDE tools that support MCP or OpenAI‑compatible APIs (Cursor, VS Code, Claude Desktop, Warp, Zed, and similar tools you’ve already documented) can route through Orq.ai to Cerebras, depending on model and integration configuration.

Check the Orq.ai integration docs for the latest supported frameworks and tools for Cerebras.

FAQs

Do I need a separate Cerebras account to use Cerebras through Orq.ai?

You can either connect your own Cerebras API key into Orq.ai (for example from the Cerebras Developer Console or Inference Cloud) or, where available, use Cerebras‑hosted models billed via Orq.ai; the exact options depend on your Orq plan, region, and how Cerebras is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Cerebras usage.

Can I route only some workflows to Cerebras and others to different providers?

Yes. You define routes per workflow in Orq.ai and decide which ones should use Cerebras vs other providers, so you can reserve Cerebras for specific performance‑sensitive or open‑model workloads while sending other tasks to different models.

Does using Cerebras through Orq.ai add latency?

Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets, while still benefiting from Cerebras’s high tokens‑per‑second capabilities.

Alternatives to

Cerebras

Anthropic

Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.

Chat

Code

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI

Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Image Generation

Reasoning

Speech

Vision

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI

Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Image Generation

Reasoning

Speech

Vision

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS

Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Reasoning

Vision

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.