Supporting image on the llm-providers/:LLM Providers page
Groq logo

Groq

on Orq.ai

Access models hosted through Groq using Orq.ai for low-latency chat, reasoning, coding, and other high-throughput AI workloads through one API.

Capabilities:

Chat

Reasoning

Code

Models Supported:

qwen/qwen3.6-27b

allam-2-7b

canopylabs/orpheus-arabic-saudi

canopylabs/orpheus-v1-english

groq/compound

groq/compound-mini

meta-llama/llama-prompt-guard-2-22m

openai/gpt-oss-safeguard-20b

qwen/qwen3-32b

openai/gpt-oss-120b

...

Provider HQ:

Mountain View, California

Access Groq through Orq.ai’s AI Router

Groq provides an inference platform built on custom Language Processing Units (LPUs), designed for extremely low latency and high throughput on open‑weight and OSS LLMs. These deployments cover high‑end reasoning, general‑purpose coding, and very cost‑efficient high‑volume use cases.

Groq models available on Orq.ai

Compare the Groq-hosted models currently available through Orq.ai, including their supported capabilities, context windows, pricing, and regional availability.

Model

Type

Context

Input / 1M

Output / 1M

qwen/qwen3.6-27b

chat

Reasoning

Vision

131K

$0.60

$3.00

allam-2-7b

chat

4.1K

$0.00

$0.00

canopylabs/orpheus-arabic-saudi

tts

Audio

4K

$40.00

$0.00

Why use Groq through Orq.ai?

Using Groq through Orq.ai lets teams add high-speed inference to a wider model stack without maintaining separate routing, evaluation, observability, and cost-control logic for each provider.

Capability

Provider

Direct

Through Orq.ai

Chat

Models hosted through Groq

Call supported models directly through Groq for chat, reasoning, coding, and other supported workloads.

Use Groq-hosted models through Orq.ai’s OpenAI-compatible endpoint while applying routing, tracing, evals, budgets, and governance controls around each request.

Code

Models hosted through Groq

Use supported models for code generation, debugging, refactoring, and agentic coding workflows.

Route coding workloads through Orq.ai, compare Groq with other providers, and monitor cost, latency, and quality from a shared control layer.

Embeddings

Supported embedding models and providers

Use supported embedding models through separate providers where required.

Route embedding workloads alongside Groq-hosted reasoning, generation, and agent workflows through the same Orq.ai platform.

Pricing

Groq pricing varies by model, usage type, provider configuration, and billing setup.

You can connect supported Groq credentials to Orq.ai or use other available access options depending on your workspace configuration. Check Orq.ai and your Groq setup for current per-model rates, quotas, and billing details.

Where supported, provider-side features such as caching or batch processing may affect effective inference costs.

Compatible frameworks and tools

Orq.ai works with OpenAI-compatible clients and common AI development frameworks. Tools that support compatible APIs or supported integration standards can also connect through Orq.ai where available.

Check the Orq.ai integration documentation for the latest setup options and compatibility information for Groq.

FAQs

Do I need a separate Groq account to use Groq through Orq.ai?

You can connect supported Groq credentials to Orq.ai or use other available access options depending on your workspace, plan, and region.

Can I route only some workflows to Groq and others to different providers?

Yes. Orq.ai lets you route different workloads to different providers, so you can use Groq where latency, throughput, or model availability fit the workload while routing other requests elsewhere.

Does using Groq through Orq.ai add latency?

Orq.ai adds a routing layer between your application and the model provider. Teams can monitor end-to-end latency and use routing, caching, and provider controls where appropriate to manage performance.

Alternatives to

Groq

Anthropic logo
Anthropic

Access Anthropic’s Claude models through Orq.ai for reasoning, coding, vision, and agentic workflows through one API.

Chat

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI logo
Open AI

Access OpenAI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.

Chat

Reasoning

Vision

Code

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI logo
Google AI

Access Google’s AI models through Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.

Chat

Reasoning

Vision

Code

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS logo
AWS

Access foundation models available through Amazon Bedrock using Orq.ai for reasoning, coding, multimodal, and other AI workloads through one API.

Chat

Reasoning

Vision

Code

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes