
Groq
on Orq.ai
Access models hosted through Groq using Orq.ai for low-latency chat, reasoning, coding, and other high-throughput AI workloads through one API.
Capabilities:
Chat
Reasoning
Code
Models Supported:
qwen/qwen3.6-27b
allam-2-7b
canopylabs/orpheus-arabic-saudi
canopylabs/orpheus-v1-english
groq/compound
groq/compound-mini
meta-llama/llama-prompt-guard-2-22m
openai/gpt-oss-safeguard-20b
qwen/qwen3-32b
openai/gpt-oss-120b
...
Provider HQ:
Mountain View, California
Access Groq through Orq.ai’s AI Router
Groq provides an inference platform built on custom Language Processing Units (LPUs), designed for extremely low latency and high throughput on open‑weight and OSS LLMs. These deployments cover high‑end reasoning, general‑purpose coding, and very cost‑efficient high‑volume use cases.
Groq models available on Orq.ai
Compare the Groq-hosted models currently available through Orq.ai, including their supported capabilities, context windows, pricing, and regional availability.
Model
Type
Context
Input / 1M
Output / 1M
qwen/qwen3.6-27b
chat
Reasoning
Vision
131K
$0.60
$3.00
allam-2-7b
chat
4.1K
$0.00
$0.00
canopylabs/orpheus-arabic-saudi
tts
Audio
4K
$40.00
$0.00
Why use Groq through Orq.ai?
Using Groq through Orq.ai lets teams add high-speed inference to a wider model stack without maintaining separate routing, evaluation, observability, and cost-control logic for each provider.
Capability | Provider | Direct | Through Orq.ai |
|---|---|---|---|
Chat | Models hosted through Groq | Call supported models directly through Groq for chat, reasoning, coding, and other supported workloads. | Use Groq-hosted models through Orq.ai’s OpenAI-compatible endpoint while applying routing, tracing, evals, budgets, and governance controls around each request. |
Code | Models hosted through Groq | Use supported models for code generation, debugging, refactoring, and agentic coding workflows. | Route coding workloads through Orq.ai, compare Groq with other providers, and monitor cost, latency, and quality from a shared control layer. |
Embeddings | Supported embedding models and providers | Use supported embedding models through separate providers where required. | Route embedding workloads alongside Groq-hosted reasoning, generation, and agent workflows through the same Orq.ai platform. |
Pricing
Groq pricing varies by model, usage type, provider configuration, and billing setup.
You can connect supported Groq credentials to Orq.ai or use other available access options depending on your workspace configuration. Check Orq.ai and your Groq setup for current per-model rates, quotas, and billing details.
Where supported, provider-side features such as caching or batch processing may affect effective inference costs.
Compatible frameworks and tools
Orq.ai works with OpenAI-compatible clients and common AI development frameworks. Tools that support compatible APIs or supported integration standards can also connect through Orq.ai where available.
Check the Orq.ai integration documentation for the latest setup options and compatibility information for Groq.
FAQs
Do I need a separate Groq account to use Groq through Orq.ai?
You can connect supported Groq credentials to Orq.ai or use other available access options depending on your workspace, plan, and region.
Can I route only some workflows to Groq and others to different providers?
Yes. Orq.ai lets you route different workloads to different providers, so you can use Groq where latency, throughput, or model availability fit the workload while routing other requests elsewhere.
Does using Groq through Orq.ai add latency?
Orq.ai adds a routing layer between your application and the model provider. Teams can monitor end-to-end latency and use routing, caching, and provider controls where appropriate to manage performance.
Alternatives to
Groq
Anthropic
Access Anthropic’s Claude models through Orq.ai for reasoning, coding, vision, and agentic workflows through one API.
Chat
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Access OpenAI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.
Chat
Reasoning
Vision
Code
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
Google AI
Access Google’s AI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.
Chat
Reasoning
Vision
Code
Models:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
AWS
Access foundation models available through Amazon Bedrock using Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.
Chat
Reasoning
Vision
Code
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5


