Supporting image on the llm-providers/:LLM Providers page
Cerebras logo

Cerebras

on Orq.ai

Access Cerebras-hosted AI models through Orq.ai for high-throughput reasoning, coding, chat, and other inference-heavy workloads through one API.

Capabilities:

Chat

Reasoning

Vision

Code

Models Supported:

cerebras/gemma-4-31b

zai-glm-4.7

cerebras/gpt-oss-120b

...

Provider HQ:

Sunnyvale, California

Access Cerebras through Orq.ai’s AI Router

Cerebras provides an inference cloud built around wafer-scale systems designed for high-throughput model serving. Its platform is suited to workloads where fast token generation, low response times, and efficient serving are important.

Cerebras models available on Orq.ai

Compare the Cerebras-hosted models currently available through Orq.ai, including their supported capabilities, context windows, pricing, and availability.

Model

Type

Context

Input / 1M

Output / 1M

cerebras/gemma-4-31b

chat

Reasoning

Vision

131K

$0.99

$1.49

zai-glm-4.7

chat

Reasoning

131K

$2.25

$2.75

cerebras/gpt-oss-120b

chat

131K

$0.35

$0.75

Why use Cerebras through Orq.ai?

Using Cerebras through Orq.ai lets teams add high-throughput inference to a wider model stack without maintaining separate routing, evaluation, observability, and cost-control logic for each provider.

Capability

Provider

Direct

Through Orq.ai

Chat

Models served through Cerebras

Call supported models directly through the Cerebras API for chat, reasoning, coding, and other supported workloads.

Use Cerebras-hosted models through Orq.ai’s OpenAI-compatible endpoint while applying routing, tracing, evals, budgets, and governance controls around each request.

Code

Models served through Cerebras

Use supported models for code generation, debugging, refactoring, and agentic coding workflows.

Route coding workloads through Orq.ai, compare Cerebras with other providers, and monitor cost, latency, and quality from a shared control layer.

Embeddings

Supported embedding models and providers

Use supported embedding models where available through your provider configuration.

Route embedding workloads alongside reasoning, generation, and agent workflows through the same Orq.ai platform.

Using Cerebras through Orq.ai lets teams add high-throughput inference to a wider model stack without maintaining separate routing, evaluation, observability, and cost-control logic for each provider.

Pricing

Cerebras model pricing varies by model, provider configuration, and billing setup.

You can connect supported Cerebras credentials to Orq.ai or use other available access options depending on your workspace configuration. Check Orq.ai and your Cerebras setup for current per-model rates, quotas, and billing details.

Compatible frameworks and tools

Orq.ai works with OpenAI-compatible clients and common AI development frameworks. Tools that support compatible APIs or supported integration standards can also connect through Orq.ai where available.

Check the Orq.ai integration documentation for the latest setup options and compatibility information for Cerebras.

FAQs

Do I need a separate Cerebras account to use Cerebras through Orq.ai?

You can connect supported Cerebras credentials to Orq.ai or use other available access options depending on your workspace, plan, and region.

Can I route only some workflows to Cerebras and others to different providers?

Yes. Orq.ai lets you route different workloads to different providers, so you can use Cerebras for workloads where throughput, latency, or provider availability fit your requirements while routing other requests elsewhere.

Does using Cerebras through Orq.ai add latency?

Orq.ai adds a routing layer between your application and the model provider. Teams can monitor end-to-end latency and use routing, caching, and provider controls where appropriate to manage performance.

Alternatives to

Cerebras

Anthropic logo
Anthropic

Access Anthropic’s Claude models through Orq.ai for reasoning, coding, vision, and agentic workflows through one API.

Chat

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI logo
Open AI

Access OpenAI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.

Chat

Reasoning

Vision

Code

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI logo
Google AI

Access Google’s AI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.

Chat

Reasoning

Vision

Code

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS logo
AWS

Access foundation models available through Amazon Bedrock using Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.

Chat

Reasoning

Vision

Code

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes