Z.ai

on Orq.ai

Use Z.ai’s GLM‑5 family through a single Orq.ai integration. Route models such as GLM‑5.2, GLM‑5.1, GLM‑5, GLM‑5‑Turbo, GLM‑4.7, and GLM‑4.7‑FlashX via Orq’s AI Router for chat, reasoning, coding, multilingual tasks, vision, and cost‑efficient high‑volume workloads.

Capabilities:

Chat

Image Generation

Reasoning

Vision

Models Supported:

glm-5.2

glm-5.1

glm-5v-turbo

glm-4.7-flashx

glm-5-turbo

Provider HQ:

Zhipu AI, based in China with global API access.

Access Z.ai through Orq.ai’s AI Router

Z.ai (Zhipu AI) is a Tsinghua‑affiliated provider offering the GLM‑5 series of text and vision models via a developer API, with long context windows and competitive pricing. Endpoints are OpenAI‑style JSON over HTTP, covering text, vision, OCR, image/video generation, speech, and agents, so they drop into many existing SDKs with minimal changes

Z.ai models available on Orq.ai

Why use Z.ai through Orq.ai?

Capability

Provider / models

Direct

Through Orq.ai

Chat & reasoning

GLM‑5.2, GLM‑5.1, GLM‑5, GLM‑4.7

Call Z.ai’s GLM endpoints directly for multilingual chat, reasoning, and assistants

Use GLM models through Orq.ai’s router, adding central routing, tracing, evals, budgets, and governance around each request

Coding

GLM‑5.2, GLM‑5.1, GLM‑5‑Turbo

Use GLM‑5 family directly for coding agents, debugging, and agent mode workflows

Route coding workloads through Orq.ai, compare GLM models against other providers, and monitor cost, latency, and quality from one control layer

Vision & tools

GLM‑5V‑Turbo, GLM‑4.6V, GLM‑OCR, GLM‑Image

Use Z.ai’s vision, OCR, and image/video generation endpoints directly

Use Orq.ai to orchestrate which steps call Z.ai (for example GLM‑5V‑Turbo for vision + text) while other steps use alternative providers, all under unified observability

This gives teams a practical way to use Z.ai where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Plans and API access

Z.ai offers:

  • API usage: pay‑per‑token across text, vision, OCR, image, video, speech, tools, and agents, with limited‑time free cached‑input storage.

  • GLM Coding Plan subscriptions: monthly/quarterly plans for IDE‑style coding assistants, separate from raw API usage.

Key text pricing per 1M tokens (USD):

  • GLM‑5.2: $1.40 input, $4.40 output, $0.26 cached input.

  • GLM‑5.1: $1.40 input, $4.40 output, $0.26 cached input.

  • GLM‑5: $1.00 input, $3.20 output, $0.20 cached input.

  • GLM‑5‑Turbo: $1.20 input, $4.00 output, $0.24 cached input.

  • GLM‑4.7: $0.60 input, $2.20 output, $0.11 cached input.

  • GLM‑4.7‑FlashX: $0.07 input, $0.40 output, $0.01 cached input.

  • GLM‑4-32B‑0414‑128K: $0.10 input, $0.10 output.

Vision, OCR, image, video, speech, and agents:

  • GLM‑5V‑Turbo: $1.20 input, $4.00 output per 1M tokens.

  • GLM‑4.6V: $0.30 input, $0.90 output per 1M tokens.

  • GLM‑OCR: $0.03 per 1M tokens.

  • GLM‑Image: $0.015 per image; CogView‑4: $0.01 per image.

  • CogVideoX and Vidu video models: $0.2–0.4 per video depending on variant.

  • GLM‑ASR‑2512: $0.03 per 1M tokens (about $0.0024 per audio minute).

  • Agents like GLM Slide/Poster Agent: around $0.70 per 1M tokens; general translation at $3 per 1M tokens.

GLM‑5 API pricing is typically 10–15× cheaper than some frontier proprietary models for comparable reasoning tasks, especially when using GLM‑4.7 Flash for high‑volume workloads.

Compatible frameworks and tools

Orq.ai exposes Z.ai through:

  • a Z.ai provider configuration in AI Router / AI Studio, where you paste your Z.ai API key, and

  • an OpenAI‑style routing layer, since GLM endpoints use JSON payloads similar to common chat/completions formats

That means:

  • Existing backend workflows using OpenAI SDKs or OpenAI‑compatible clients can point at Orq’s base URL and route some traffic to GLM‑5 models without changing request payloads.

  • Agents, code assistants, and tools that integrate with Orq.ai can be configured so that specific steps (for example, multilingual coding or GLM‑5V‑Turbo vision) are served by Z.ai, while other steps use different providers, all sharing the same observability and governance layer.

Check the Orq.ai integration docs for the latest supported frameworks and tools for Z.ai.

FAQs

Do I need a separate Z.ai account to use Z.ai through Orq.ai?

Yes. You create a Z.ai developer account, obtain an API key, then configure it in Orq.ai’s AI Router; Orq.ai then gives you one place to manage routing, observability, and cost controls around that GLM usage.

Can I route only some workflows to Z.ai and others to different providers?

Yes. You define routes per workflow in Orq.ai and choose which ones should use GLM‑5 vs other models, so you can reserve Z.ai for multilingual coding, long‑context reasoning, or cost‑ optimized workloads while sending other tasks to frontier or regional providers.

Does using Z.ai through Orq.ai add latency?

Orq.ai is designed as a lightweight router, so the added overhead is small compared to GLM’s own inference time. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets while gaining visibility and control.

Alternatives to

Z.ai

Anthropic

Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.

Chat

Code

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI

Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Image Generation

Reasoning

Speech

Vision

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI

Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Image Generation

Reasoning

Speech

Vision

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS

Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Reasoning

Vision

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.