Groq

on Orq.ai

Use Groq’s LPU‑accelerated models through a single Orq.ai API. Route popular open models such as Llama 3.1, Llama 3.3, Qwen, Gemma, Mixtral, and GPT‑OSS variants via Orq’s AI Router for ultra‑fast chat, reasoning, and coding workloads.

Capabilities:

Chat

Reasoning

Speech

Vision

Models Supported:

qwen/qwen3.6-27b

allam-2-7b

canopylabs/orpheus-arabic-saudi

canopylabs/orpheus-v1-english

groq/compound

Provider HQ:

Groq (GroqCloud), with operations centered in the United States.

Access Groq through Orq.ai’s AI Router

Groq provides an inference platform built on custom Language Processing Units (LPUs), designed for extremely low latency and high throughput on open‑weight and OSS LLMs. These deployments cover high‑end reasoning, general‑purpose coding, and very cost‑efficient high‑volume use cases.

Orq.ai supports major GroqCloud models (for example Llama 3.1 8B/70B, Llama 3.3 70B, GPT‑OSS‑class models, Qwen3‑32B, Gemma 2, Mixtral), with availability depending on which models you enable in the Groq provider and your workspace configuration.

Groq models available on Orq.ai

Model (example)

Type

Context

Best for

Pricing tier (reference)

Llama 3.3 70B Instruct / GPT‑OSS‑120B on Groq

Chat / reasoning / coding

Around 130K tokens (check Groq / Orq docs for current limits)

Complex reasoning, analysis, coding, and high‑stakes tasks where you want large open models served at very high speed

Premium – Llama 3.3 70B around 0.59 USD / 1M input tokens and 0.79 USD / 1M output tokens; GPT‑OSS‑120B around 0.15 USD / 1M input and 0.60–0.75 USD / 1M output

Qwen3‑32B, Mixtral 8x7B, Gemma 2 9B

Chat / reasoning / coding

Context windows typically 40K–130K tokens

Everyday production workloads, RAG, coding, and product features balancing quality with cost and latency

Mid‑tier – Qwen3‑32B around 0.29 USD / 1M input tokens and 0.59 USD / 1M output; Mixtral and Gemma 2 9B in the 0.20–0.34 USD per 1M token range

Llama 3.1 8B Instruct / smaller guard models

Chat / fast / cost‑efficient

Around 128K context for 8B models; 512 for guard models

Fast, low‑cost tasks such as classification, extraction, lightweight chat, routing, and high‑volume support workflows

Cost‑efficient – Llama 3.1 8B around 0.05 USD / 1M input tokens and 0.08 USD / 1M output; prompt‑guard models as low as 0.03–0.04 USD input and output per 1M tokens

Pricing tiers here are approximate and based on Groq’s public API pricing; Orq.ai may apply its own billing or BYOK mapping. Always check Orq.ai’s pricing page and your configured Groq provider for current per‑model rates, quotas, and billing details.

Why use Groq through Orq.ai?

Capability

Provider / models

Direct

Through Orq.ai

Chat

GroqCloud chat models (for example Llama 3.1/3.3, GPT‑OSS, Qwen3)

Call GroqCloud directly via its OpenAI‑style API for chat, reasoning, coding, and RAG

Use Groq models through Orq.ai’s OpenAI‑compatible endpoint, adding routing, tracing, evals, budgets, and governance controls around each request

Code

Llama, Qwen, and GPT‑OSS models tuned or suitable for coding

Use Groq directly for code generation, debugging, refactoring, and agentic coding workflows at very low latency

Route coding workloads through Orq.ai, compare Groq‑hosted models against other providers, and monitor cost, latency, and quality from one control layer

Embeddings

Embedding providers configured in Orq.ai

Groq focuses on high‑speed LLM inference; embeddings may be handled by separate models or providers

Use Orq.ai to route embedding workloads to supported embedding providers while keeping Groq for reasoning, generation, and agent steps.

This gives teams a practical way to use Grok where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Model rates

Groq model pricing may differ depending on whether you:

  • connect your own GroqCloud account and API key (BYOK), or

  • use Groq models billed through Orq.ai where available.console.groq+1

Current public data shows:

  • Entry‑tier models like Llama 3.1 8B around 0.05 USD / 1M input and 0.08 USD / 1M output tokens

  • Mid‑tier models such as Qwen3‑32B or Mixtral in the 0.29–0.34 USD per 1M token range

  • Large models like Llama 3.3 70B around 0.59 USD / 1M input and 0.79 USD / 1M output tokens

Groq also offers discounts via prompt caching and batch APIs for non‑real‑time workloads.

Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details.

Compatible frameworks and tools

Orq.ai exposes Groq models through:

  • a Groq provider configuration pointing at Groq’s OpenAI‑style endpoint, and

  • an OpenAI‑compatible API layer for applications that expect that interface

That means:

  • Popular AI frameworks (for example, OpenAI‑compatible clients, AI SDKs, and orchestration libraries) can talk to Groq via Orq’s router simply by pointing to Orq’s OpenAI‑compatible endpoint.console.

  • Code assistants and IDE tools that support MCP or OpenAI‑compatible APIs (Cursor, VS Code, Claude Desktop, Warp, Zed, and similar tools you’ve documented) can route through Orq.ai to Groq, depending on model and integration configuration.

Check the Orq.ai integration docs for the latest supported frameworks and tools for Groq.

FAQs

Do I need a separate Groq account to use Groq through Orq.ai?

You can either connect your own Groq API key into Orq.ai or, where available, use Groq models billed via Orq.ai; the exact options depend on your Orq plan, region, and how Groq is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Groq usage.

Can I route only some workflows to Groq and others to different providers?

Yes. You define routes per workflow in Orq.ai and decide which ones should use Groq vs other models, so you can reserve Groq for latency‑sensitive, high‑volume, or specific open‑model workloads while sending other tasks to different providers.

Does using Groq through Orq.ai add latency?

Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to the model’s own latency. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets while still benefiting from Groq’s LPU‑level speed.

Alternatives to

Groq

Anthropic

Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.

Chat

Code

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI

Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Image Generation

Reasoning

Speech

Vision

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI

Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Image Generation

Reasoning

Speech

Vision

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS

Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Reasoning

Vision

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.