
Microsoft Azure
on Orq.ai
Use Azure‑hosted OpenAI models through a single Orq.ai API. Route models such as GPT‑4o, GPT‑4o‑mini, and other Azure OpenAI deployments via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Capabilities:
Chat
Reasoning
Vision
Models Supported:
gpt-5.6-luna
gpt-5.6-luna
gpt-5.6-sol
gpt-5.6-sol
gpt-5.6-terra
Provider HQ:
Microsoft, headquartered in Redmond, Washington, with Azure AI / Azure OpenAI regions across North America, Europe, and other geographies.
Access Azure through Orq.ai’s AI Router
Azure OpenAI (now part of Azure AI Foundry) is Microsoft’s managed platform for OpenAI foundation models, exposing GPT‑4‑class and GPT‑3.5‑class models (and newer generations) under a single Azure service with enterprise controls. These models cover high‑end reasoning, general‑purpose coding, and cost‑efficient high‑volume use cases.
Orq.ai supports major Azure OpenAI deployments (for example GPT‑4o, GPT‑4o‑mini, GPT‑4.1 and successors), with availability depending on region, deployment configuration, and your workspace setup.
Azure models available on Orq.ai
Pricing tiers here are approximate and based on Alibaba’s public Qwen API pricing; Orq.ai may apply its own billing or BYOK mapping. Always check Orq.ai’s pricing page and your configured provider for current per‑model rates, quotas, and billing details.
Model | Type | Context / scope | Best for | Pricing tier |
|---|---|---|---|---|
GPT‑4o (Azure deployment) | Chat / reasoning / vision | Large context (check Azure / Orq docs for current limits) | Complex reasoning, analysis, coding, multimodal tasks, and high‑stakes workloads where you want Microsoft’s strongest generally‑available OpenAI model on Azure | Premium – Azure OpenAI standard API pricing; typical range around 2–3 USD / 1M input tokens and 8–10 USD / 1M output tokens, depending on region and plan |
GPT‑4o‑mini (Azure deployment) | Chat / reasoning / coding | Medium–large context (verify in Azure / Orq docs) | Everyday production workloads, RAG, coding, product features, and workflows that balance quality with cost and latency | Mid‑tier – Azure OpenAI pricing; often around 0.1–0.2 USD / 1M input tokens and 0.5–0.6 USD / 1M output tokens, depending on region and caching setu |
GPT‑3.5 / smaller Azure OpenAI models | Chat / fast / cost‑efficient | Medium context (confirm in docs) | Fast, lower‑cost tasks such as classification, extraction, lightweight chat, routing, and high‑volume support workflows | Cost‑efficient – Azure OpenAI pricing; older GPT‑3.5‑class tiers can be as low as 0.5 USD / 1M input tokens and 1.5 USD / 1M output tokens |
Why use Azure through Orq.ai?
Capability | Provider / models | Direct | Through Orq.ai |
|---|---|---|---|
Chat | Qwen‑Max, Qwen‑Plus, Qwen‑Turbo and related Qwen3.x models | Call Qwen directly via Alibaba Cloud Model Studio’s API for chat, reasoning, coding, and multimodal tasks | Use Qwen through Orq.ai’s OpenAI‑compatible endpoint, adding routing, tracing, evals, budgets, and governance controls around each request |
Code | Qwen‑Max / Plus / Turbo (and newer Qwen3.x tiers tuned for coding) | Use Qwen directly for code generation, debugging, refactoring, and agentic coding workflows | Route coding workloads through Orq.ai, compare Qwen against other providers, and monitor cost, latency, and quality from one control layer |
Embeddings | Embedding providers configured in Orq.ai | Qwen models focus on chat/reasoning; embeddings are typically handled by separate models or providers | Use Orq.ai to route embedding workloads to supported embedding providers while keeping Qwen for reasoning, generation, and agent steps |
This gives teams a practical way to use Azure‑hosted OpenAI models where they perform best while centralising routing, observability, evals, and cost controls across the wider model stack.
Pricing
Azure OpenAI model pricing may differ depending on whether you:
connect your own Azure AI / Azure OpenAI account and deployments (BYOK), or
use Azure‑hosted models billed through Orq.ai where available.
Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details.
Compatible frameworks and tools
Orq.ai exposes Azure‑hosted OpenAI models through:
an OpenAI‑compatible API layer, and
native SDKs and router integrations where applicable
That means:
Popular AI frameworks (for example, OpenAI‑compatible clients and orchestration libraries) can talk to Azure OpenAI deployments via Orq’s router.
Code assistants and IDE tools that support MCP or OpenAI‑compatible APIs (Cursor, VS Code, Claude Desktop, Warp, Zed, and similar tools you’ve documented) can route through Orq.ai to Azure, depending on model and integration configuration.
Check the Orq.ai integration docs for the latest supported frameworks and tools for Azure OpenAI / Azure AI Foundry.
FAQs
Do I need a separate Azure account to use Azure OpenAI through Orq.ai?
You can either connect your own Azure AI / Azure OpenAI account and credentials into Orq.ai or, where available, use Azure‑hosted models billed via Orq.ai; the exact options depend on your Orq plan, region, and how Azure is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Azure model usage.
Can I route only some workflows to Azure and others to different providers?
Yes. You define routes per workflow in Orq.ai and decide which ones should use Azure‑hosted models vs other providers, so you can reserve Azure for specific regions, compliance needs, or workloads while sending other tasks to different models.
Does using Azure through Orq.ai add latency?
Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets.
Alternatives to
Microsoft Azure
Anthropic
Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.
Chat
Code
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Image Generation
Reasoning
Speech
Vision
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
Google AI
Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Image Generation
Reasoning
Speech
Vision
Models:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
AWS
Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Reasoning
Vision
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5


