
Moonshot AI
on Orq.ai
Use Moonshot AI’s Kimi models through a single Orq.ai API. Route models such as Kimi K2.6, Kimi K2.5, and Moonshot‑v1‑8k/32k/128k via Orq’s AI Router for chat, reasoning, coding, and long‑context workloads.
Capabilities:
Chat
Reasoning
Vision
Models Supported:
Kimi K3
kimi-k2.7-code-highspeed
moonshot-v1-128k
moonshot-v1-128k-vision-preview
moonshot-v1-32k
Provider HQ:
Moonshot AI (Kimi), based in China with global developer access via the Kimi API platform.
Access Moonshot AI through Orq.ai’s AI Router
Moonshot AI (Kimi) is a leading Chinese LLM provider offering the Kimi family of long‑context language models via a managed cloud API, with a consumer app and separate token‑based API platform. These models cover high‑end reasoning, code, and deep long‑context analysis, with context windows up to around 256K-300K tokens.
Moonshot AI models available on Orq.ai
Model | Type | Context / scope | Best for | Pricing tier |
|---|---|---|---|---|
Kimi K2.6 (flagship) | Chat / reasoning / coding / long context | Up to around 200K–256K tokens (check Kimi / Orq docs) | Complex reasoning, deep document analysis, coding, and high‑stakes tasks where you want Moonshot’s strongest Kimi model | Premium – K2.6 API pricing around 0.95 USD / 1M input tokens and 4.00 USD / 1M output tokens on the latest public data.developer |
Kimi K2.5 (reasoning) | Chat / reasoning / long context | Around 256K context (verify in docs) | Long‑context RAG, analytics, and production workloads that balance quality with cost and latency | Mid‑tier – K2.5 direct API pricing around 0.60 USD / 1M input tokens and 2.50–3.00 USD / 1M output tokens; some aggregators list similar rates |
Moonshot‑v1‑8k / 32k / 128k | Chat / fast / cost‑efficient | 8K / 32K / 128K tokens depending on tier | Fast, lower‑cost tasks such as classification, extraction, lightweight chat, routing, and smaller long‑context jobs | Cost‑efficient – examples: Moonshot‑v1‑8k at about 0.20 USD / 1M input tokens; Moonshot‑v1‑32k around 1.00 USD / 1M input; Moonshot‑v1‑128k around 2.00 USD / 1M input tokens, with separate output pricing where published |
Why use Moonshot AI through Orq.ai?
Capability | Provider / models | Direct | Through Orq.ai |
|---|---|---|---|
Chat | Kimi and Moonshot‑v1 chat models (for example K2.6, K2.5, K2, Moonshot‑v1‑8k/32k/128k) | Call Kimi / Moonshot models directly via the Kimi API platform for chat, reasoning, coding, and long‑context tasks | Use Moonshot models through Orq.ai’s OpenAI‑compatible endpoint, adding routing, tracing, evals, budgets, and governance controls around each request |
Code | Kimi models suitable for coding/analysis | Use Kimi directly for code generation, debugging, refactoring, and agentic coding workflows, especially in Chinese‑language and APAC settings | Route coding workloads through Orq.ai, compare Moonshot models against other providers, and monitor cost, latency, and quality from one control layer |
Long‑context analysis | K2.6 / K2.5 / high‑context Moonshot‑v1 tiers | Use the Kimi platform directly for 128K–256K‑token context use cases, such as multi‑document analysis and long‑form chat | Use Orq.ai to route only the workflows that truly need Kimi’s long context, while keeping shorter tasks on cheaper providers, with unified observability and controls |
This gives teams a practical way to use Moonshot AI where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.
Pricing
Plans and API access
Moonshot AI has two main surfaces:
the Kimi app (consumer and pro memberships), and
the Kimi / Moonshot API (token‑based billing).platform
For API usage (2026 public data):
Entry‑tier / cheaper models:
Moonshot‑v1‑8k: about 0.20 USD / 1M input tokens
Some Moonshot models on third‑party gateways start around 0.39–0.60 USD / 1M input tokens
Tier / model | Input / price | Output / included | Cached / notes |
|---|---|---|---|
Mid‑tier: | * Kimi K2.5: around 0.60 USD / 1M input tokens and 2.50–3.00 USD / 1M output tokens, with a context window around 256K tokens. | ||
Premium: | * Kimi K2.6: about 0.95 USD / 1M input tokens and 4.00 USD / 1M output tokens as Moonshot’s flagship reasoning model | Moonshot also offers: | * a free tier for the Kimi app (roughly 30–50 messages/day reported) and * paid memberships (for example “Moderato” at about 19 USD/month) that expand app‑side quotas and features, separate from raw API usage |
Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details.
Compatible frameworks and tools
Orq.ai exposes Moonshot AI through:
a Moonshot provider configuration in AI Gateway / Model Garden, where you paste your Moonshot/Kimi API key, and
an OpenAI‑compatible API layer for applications that expect that interface
That means:
Existing backend workflows that already talk to Orq’s OpenAI‑compatible endpoint can be wired to call Kimi / Moonshot models through Orq.ai without a separate direct integration
Agents, code assistants, and tools that integrate with Orq.ai (for example, via OpenAI‑compatible or HTTP tools) can be configured so that long‑context or Chinese‑language steps are served by Kimi, while other steps use different providers
Check the Orq.ai integration docs for the latest supported frameworks and tools for Moonshot AI.
FAQs
Do I need a separate Moonshot AI / Kimi account to use Moonshot through Orq.ai?
You can either connect your own Moonshot / Kimi API key from platform.kimi.ai into Orq.ai or, where available, use Moonshot usage billed via Orq.ai; the exact options depend on your Orq plan, region, and how Moonshot is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Moonshot usage.
Can I route only some workflows to Moonshot and others to different providers?
Yes. You define routes per workflow in Orq.ai and decide which ones should use Kimi / Moonshot vs other models, so you can reserve Moonshot for long‑context, Chinese‑language, or specific reasoning workloads while sending other tasks to different providers.
Does using Moonshot AI through Orq.ai add latency?
Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets.
Alternatives to
Moonshot AI
Anthropic
Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.
Chat
Code
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Image Generation
Reasoning
Speech
Vision
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
Google AI
Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Image Generation
Reasoning
Speech
Vision
Models:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
AWS
Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Reasoning
Vision
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5


