
Wafer
on Orq.ai
Access models hosted through Wafer using Orq.ai for reasoning, coding, agentic, and high-throughput AI workloads through one API.
Capabilities:
Chat
Reasoning
Code
Models Supported:
GLM-5.2
GLM5.2-Fast
GLM-5.1
Kimi-K2.6
Qwen3.5-397B-A17B
Qwen3.6-35B-A3B
...
Provider HQ:
San Francisco, California
Access Wafer through Orq.ai’s AI Router
Wafer is an AI inference platform focused on serving open and open-weight models with high-throughput, predictable performance. Its platform supports language, reasoning, coding, and agentic workloads through managed inference and compatible API interfaces.
Wafer models available on Orq.ai
Compare the Wafer-hosted models currently available through Orq.ai, including their supported capabilities, context windows, pricing, and availability.
Model
Type
Context
Input / 1M
Output / 1M
GLM-5.2
chat
Reasoning
1M
$1.20
$4.10
GLM5.2-Fast
chat
Reasoning
1M
$3.00
$10.25
GLM-5.1
chat
Reasoning
203K
$1.00
$3.20
Why use Wafer through Orq.ai?
Using Wafer through Orq.ai lets teams combine high-throughput model inference with a wider model stack without maintaining separate routing, evaluation, observability, and cost-control logic for each provider.
Capability | Provider | Direct | Through Orq.ai |
|---|---|---|---|
Chat & reasoning | Models hosted through Wafer | Call supported models directly through Wafer for chat, reasoning, generation, and other language-model workloads. | Route Wafer-hosted models through Orq.ai while applying routing, tracing, evals, budgets, and governance controls around each request. |
Coding | Models hosted through Wafer | Use supported models for code generation, debugging, refactoring, and agentic coding workflows. | Route coding workloads through Orq.ai, compare Wafer with other providers, and monitor cost, latency, and quality from a shared control layer. |
High-throughput agents | Supported Wafer models | Use Wafer directly for workloads that require frequent or sustained model inference. | Route throughput-heavy agent workloads to Wafer while using other providers for tasks with different capability, cost, or regional requirements. |
This gives teams a practical way to use Wafer where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.
Pricing
Wafer pricing varies by model, usage volume, access method, and billing configuration.
Depending on the service being used, pricing may be consumption-based or tied to a subscription or usage allowance. Provider-side caching, privacy controls, and other service options may also affect how workloads are billed.
You can connect supported Wafer credentials to Orq.ai or use other available access options depending on your workspace configuration. Check Orq.ai and your Wafer setup for current rates, quotas, and billing details.
Compatible frameworks and tools
Wafer supports commonly used API formats for model inference, while Orq.ai provides a shared routing layer for supported integrations.
Existing applications using compatible clients can incorporate Wafer-hosted models into Orq.ai workflows depending on the model, endpoint, and workspace configuration.
Check the Orq.ai integration documentation for the latest setup options and compatibility information for Wafer.
FAQs
Do I need a separate Wafer account to use Wafer through Orq.ai?
You can connect supported Wafer credentials to Orq.ai or use other available access options depending on your workspace, plan, and region.
Can I route only some workflows to Wafer and others to different providers?
Yes. Orq.ai lets you route different workloads to different providers, so you can use Wafer for high-throughput, coding, agentic, or privacy-sensitive workloads while routing other requests elsewhere.
Does using Wafer through Orq.ai add latency?
Orq.ai adds a routing layer between your application and the model provider. Teams can monitor end-to-end latency and use routing, caching, and provider controls where appropriate to manage performance.
Alternatives to
Wafer
Anthropic
Access Anthropic’s Claude models through Orq.ai for reasoning, coding, vision, and agentic workflows through one API.
Chat
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Access OpenAI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.
Chat
Reasoning
Vision
Code
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
Google AI
Access Google’s AI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.
Chat
Reasoning
Vision
Code
Models:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
AWS
Access foundation models available through Amazon Bedrock using Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.
Chat
Reasoning
Vision
Code
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5


