Supporting image on the llm-providers/:LLM Providers page
Wafer logo

Wafer

on Orq.ai

Access models hosted through Wafer using Orq.ai for reasoning, coding, agentic, and high-throughput AI workloads through one API.

Capabilities:

Chat

Reasoning

Code

Models Supported:

GLM-5.2

GLM5.2-Fast

GLM-5.1

Kimi-K2.6

Qwen3.5-397B-A17B

Qwen3.6-35B-A3B

...

Provider HQ:

San Francisco, California

Access Wafer through Orq.ai’s AI Router

Wafer is an AI inference platform focused on serving open and open-weight models with high-throughput, predictable performance. Its platform supports language, reasoning, coding, and agentic workloads through managed inference and compatible API interfaces.

Wafer models available on Orq.ai

Compare the Wafer-hosted models currently available through Orq.ai, including their supported capabilities, context windows, pricing, and availability.

Model

Type

Context

Input / 1M

Output / 1M

GLM-5.2

chat

Reasoning

1M

$1.20

$4.10

GLM5.2-Fast

chat

Reasoning

1M

$3.00

$10.25

GLM-5.1

chat

Reasoning

203K

$1.00

$3.20

Why use Wafer through Orq.ai?

Using Wafer through Orq.ai lets teams combine high-throughput model inference with a wider model stack without maintaining separate routing, evaluation, observability, and cost-control logic for each provider.

Capability

Provider

Direct

Through Orq.ai

Chat & reasoning

Models hosted through Wafer

Call supported models directly through Wafer for chat, reasoning, generation, and other language-model workloads.

Route Wafer-hosted models through Orq.ai while applying routing, tracing, evals, budgets, and governance controls around each request.

Coding

Models hosted through Wafer

Use supported models for code generation, debugging, refactoring, and agentic coding workflows.

Route coding workloads through Orq.ai, compare Wafer with other providers, and monitor cost, latency, and quality from a shared control layer.

High-throughput agents

Supported Wafer models

Use Wafer directly for workloads that require frequent or sustained model inference.

Route throughput-heavy agent workloads to Wafer while using other providers for tasks with different capability, cost, or regional requirements.

This gives teams a practical way to use Wafer where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Wafer pricing varies by model, usage volume, access method, and billing configuration.

Depending on the service being used, pricing may be consumption-based or tied to a subscription or usage allowance. Provider-side caching, privacy controls, and other service options may also affect how workloads are billed.

You can connect supported Wafer credentials to Orq.ai or use other available access options depending on your workspace configuration. Check Orq.ai and your Wafer setup for current rates, quotas, and billing details.

Compatible frameworks and tools

Wafer supports commonly used API formats for model inference, while Orq.ai provides a shared routing layer for supported integrations.

Existing applications using compatible clients can incorporate Wafer-hosted models into Orq.ai workflows depending on the model, endpoint, and workspace configuration.

Check the Orq.ai integration documentation for the latest setup options and compatibility information for Wafer.

FAQs

Do I need a separate Wafer account to use Wafer through Orq.ai?

You can connect supported Wafer credentials to Orq.ai or use other available access options depending on your workspace, plan, and region.

Can I route only some workflows to Wafer and others to different providers?

Yes. Orq.ai lets you route different workloads to different providers, so you can use Wafer for high-throughput, coding, agentic, or privacy-sensitive workloads while routing other requests elsewhere.

Does using Wafer through Orq.ai add latency?

Orq.ai adds a routing layer between your application and the model provider. Teams can monitor end-to-end latency and use routing, caching, and provider controls where appropriate to manage performance.

Alternatives to

Wafer

Anthropic logo
Anthropic

Access Anthropic’s Claude models through Orq.ai for reasoning, coding, vision, and agentic workflows through one API.

Chat

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI logo
Open AI

Access OpenAI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.

Chat

Reasoning

Vision

Code

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI logo
Google AI

Access Google’s AI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.

Chat

Reasoning

Vision

Code

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS logo
AWS

Access foundation models available through Amazon Bedrock using Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.

Chat

Reasoning

Vision

Code

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes