
Together AI
on Orq.ai
Use Together AI’s open‑model platform through a single Orq.ai integration. Route leading models such as Llama 4 Scout/Maverick, Llama 3.1/3.3, DeepSeek V3/R1, Mixtral, Qwen, Gemma, and many code‑tuned variants via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Capabilities:
Chat
Reasoning
Vision
Models Supported:
GLM-5.2
Kimi K2.7 Code
Kimi K2.6
DeepSeek V3
meta-llama/Llama-3.3-70B-Instruct-Turbo
Provider HQ:
Together AI, a US‑based infrastructure and model‑hosting provider.
Access Together AI through Orq.ai’s AI Router
Together AI hosts 200+ open‑source and open‑weight models behind a unified, serverless API, with pay‑per‑token pricing across text, code, image, and audio. Its endpoints are OpenAI‑style, so they drop into existing OpenAI SDKs while giving access to highly cost‑efficient Llama, DeepSeek, Qwen, Mixtral, and other families.
Together AI models available on Orq.ai
Model | Type | Context / scope | Best for | Pricing tier |
|---|---|---|---|---|
Llama 4 Maverick (serverless) | Chat / reasoning / coding | Up to ~1M context | Complex reasoning, code generation, and general chat where you want a strong, long‑context Llama 4 model | Premium – around $0.20 / 1M input tokens and $0.60 / 1M output tokens, with large 10M context variants listed in some catalogs |
Llama 3.3 / 3.1 70B Instruct Turbo | Chat / reasoning / RAG | 128K context | Everyday production workloads, RAG, assistants, and product features with high quality but lower cost than closed models | Mid‑tier – typical rates around $0.88 / 1M input and $0.88 / 1M output for 70B Llama 3.1/3.3 |
Llama 4 Scout, Llama 3.1 8B, DeepSeek V3 | Chat / fast / cost‑efficient | 128K–10M context (per model) | Fast, lower‑cost tasks such as high‑volume chat, summarisation, classification, and routing, with strong price/performance | Cost‑efficient – examples: Llama 4 Scout at $0.11 input / $0.34 output, Llama 3.1 8B at $0.18 / $0.18, DeepSeek V3 around $0.27 input / $1.10 output depending on catalog |
Why use Together AI through Orq.ai?
Capability | Provider / models | Direct | Through Orq.ai |
|---|---|---|---|
Chat & reasoning | Llama 4 Scout/Maverick, Llama 3.x, Mixtral, DeepSeek V3/R1 | Call Together’s serverless models directly for chat, reasoning, and assistants | Use Together models through Orq.ai’s router, adding central routing, tracing, evals, budgets, and governance around each request |
Coding | DeepSeek V3/R1, DeepSeek Coder, Llama 4 Maverick, Qwen 2.5 | Use Together directly for code generation, debugging, and code‑aware assistants | Route coding workloads through Orq.ai, compare Together‑served models against other providers, and monitor cost, latency, and quality from one control layer |
Multimodal | Llama vision variants, image/video models in Together catalog | Use Together’s image/video/audio endpoints directly for multimodal workflows | Use Orq.ai to orchestrate which steps call Together (for example, long‑context Llama 4 or DeepSeek vision) while other steps use alternative providers, all under unified observability |
This gives teams a practical way to use Together where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.
Pricing
Plans and API access
Together AI offers:
Serverless inference: pay‑per‑token for text/code models with no minimums or provisioning
Dedicated endpoints and GPU clusters: custom hosting and fine‑tuning, billed separately
For serverless text/code models (2026 public data):
Tiny / budget models (for example Llama 4 Scout): prices start around $0.11 / 1M input tokens and $0.34 / 1M output tokens.
Small / 8B class models: Llama 3.1/3.3 8B often around $0.18 / 1M input and $0.18 / 1M output.
Medium / 70B models: Llama 3.1/3.3 70B around $0.88 / 1M input and $0.88 / 1M output.
Larger or reasoning‑heavy models: DeepSeek R1 and Llama 3.1 405B are priced higher (for example DeepSeek R1 at $3.00 input / $7.00 output, Llama 3.1 405B at $3.50 / 1M per side).
A notable pattern is that many Together models have identical input and output pricing, which can dramatically reduce costs versus providers that charge much more for output.
Check Together’s pricing page and your Orq.ai provider configuration for current rates and any free‑credit offers when using Together via Orq.ai
Compatible frameworks and tools
Orq.ai exposes Together AI through:
a Together AI provider configuration in AI Router, where you add your Together API key, and
an OpenAI‑style integration, since Together’s endpoints are compatible with the common chat/completions format
That means:
Existing backend workflows using OpenAI SDKs or OpenAI‑compatible clients can switch their base URL to Orq.ai and route some traffic to Together models without changing request payloads.
Agents, code assistants, and tools that integrate with Orq.ai can be configured so that specific steps (for example, long‑context Llama 4 summarisation or DeepSeek reasoning) are served by Together, while other steps use different providers, all sharing the same observability and governance layer.
Check the Orq.ai integration docs for the latest supported frameworks and tools for Together AI.
FAQs
Do I need a separate Together AI account to use Together through Orq.ai?
Yes. You sign up with Together AI, obtain an API key, then configure it in Orq.ai’s AI Router; Orq.ai then gives you one place to manage routing, observability, and cost controls around that Together usage.
Can I route only some workflows to Together AI and others to different providers?
Yes. You define routes per workflow in Orq.ai and choose which ones should use Together vs other providers, so you can reserve Together for open‑model, cost‑sensitive, or long‑context workloads while sending other tasks to proprietary or regional providers.
Does using Together AI through Orq.ai add latency?
Orq.ai is designed as a lightweight router, so the added overhead is small compared to model inference time. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets while gaining visibility and control.
Alternatives to
Together AI
Anthropic
Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.
Chat
Code
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Image Generation
Reasoning
Speech
Vision
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
Google AI
Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Image Generation
Reasoning
Speech
Vision
Models:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
AWS
Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Reasoning
Vision
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5


