Together AI

on Orq.ai

Use Together AI’s open‑model platform through a single Orq.ai integration. Route leading models such as Llama 4 Scout/Maverick, Llama 3.1/3.3, DeepSeek V3/R1, Mixtral, Qwen, Gemma, and many code‑tuned variants via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Capabilities:

Chat

Reasoning

Vision

Models Supported:

GLM-5.2

Kimi K2.7 Code

Kimi K2.6

DeepSeek V3

meta-llama/Llama-3.3-70B-Instruct-Turbo

Provider HQ:

Together AI, a US‑based infrastructure and model‑hosting provider.

Access Together AI through Orq.ai’s AI Router

Together AI hosts 200+ open‑source and open‑weight models behind a unified, serverless API, with pay‑per‑token pricing across text, code, image, and audio. Its endpoints are OpenAI‑style, so they drop into existing OpenAI SDKs while giving access to highly cost‑efficient Llama, DeepSeek, Qwen, Mixtral, and other families.

Together AI models available on Orq.ai

Model

Type

Context / scope

Best for

Pricing tier

Llama 4 Maverick (serverless)

Chat / reasoning / coding

Up to ~1M context

Complex reasoning, code generation, and general chat where you want a strong, long‑context Llama 4 model

Premium – around $0.20 / 1M input tokens and $0.60 / 1M output tokens, with large 10M context variants listed in some catalogs

Llama 3.3 / 3.1 70B Instruct Turbo

Chat / reasoning / RAG

128K context

Everyday production workloads, RAG, assistants, and product features with high quality but lower cost than closed models

Mid‑tier – typical rates around $0.88 / 1M input and $0.88 / 1M output for 70B Llama 3.1/3.3

Llama 4 Scout, Llama 3.1 8B, DeepSeek V3

Chat / fast / cost‑efficient

128K–10M context (per model)

Fast, lower‑cost tasks such as high‑volume chat, summarisation, classification, and routing, with strong price/performance

Cost‑efficient – examples: Llama 4 Scout at $0.11 input / $0.34 output, Llama 3.1 8B at $0.18 / $0.18, DeepSeek V3 around $0.27 input / $1.10 output depending on catalog

Why use Together AI through Orq.ai?

Capability

Provider / models

Direct

Through Orq.ai

Chat & reasoning

Llama 4 Scout/Maverick, Llama 3.x, Mixtral, DeepSeek V3/R1

Call Together’s serverless models directly for chat, reasoning, and assistants

Use Together models through Orq.ai’s router, adding central routing, tracing, evals, budgets, and governance around each request

Coding

DeepSeek V3/R1, DeepSeek Coder, Llama 4 Maverick, Qwen 2.5

Use Together directly for code generation, debugging, and code‑aware assistants

Route coding workloads through Orq.ai, compare Together‑served models against other providers, and monitor cost, latency, and quality from one control layer

Multimodal

Llama vision variants, image/video models in Together catalog

Use Together’s image/video/audio endpoints directly for multimodal workflows

Use Orq.ai to orchestrate which steps call Together (for example, long‑context Llama 4 or DeepSeek vision) while other steps use alternative providers, all under unified observability

This gives teams a practical way to use Together where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Plans and API access

Together AI offers:

  • Serverless inference: pay‑per‑token for text/code models with no minimums or provisioning

  • Dedicated endpoints and GPU clusters: custom hosting and fine‑tuning, billed separately

For serverless text/code models (2026 public data):

  • Tiny / budget models (for example Llama 4 Scout): prices start around $0.11 / 1M input tokens and $0.34 / 1M output tokens.

  • Small / 8B class models: Llama 3.1/3.3 8B often around $0.18 / 1M input and $0.18 / 1M output.

  • Medium / 70B models: Llama 3.1/3.3 70B around $0.88 / 1M input and $0.88 / 1M output.

  • Larger or reasoning‑heavy models: DeepSeek R1 and Llama 3.1 405B are priced higher (for example DeepSeek R1 at $3.00 input / $7.00 output, Llama 3.1 405B at $3.50 / 1M per side).

A notable pattern is that many Together models have identical input and output pricing, which can dramatically reduce costs versus providers that charge much more for output.

Check Together’s pricing page and your Orq.ai provider configuration for current rates and any free‑credit offers when using Together via Orq.ai

Compatible frameworks and tools

Orq.ai exposes Together AI through:

  • a Together AI provider configuration in AI Router, where you add your Together API key, and

  • an OpenAI‑style integration, since Together’s endpoints are compatible with the common chat/completions format

That means:

  • Existing backend workflows using OpenAI SDKs or OpenAI‑compatible clients can switch their base URL to Orq.ai and route some traffic to Together models without changing request payloads.

  • Agents, code assistants, and tools that integrate with Orq.ai can be configured so that specific steps (for example, long‑context Llama 4 summarisation or DeepSeek reasoning) are served by Together, while other steps use different providers, all sharing the same observability and governance layer.

Check the Orq.ai integration docs for the latest supported frameworks and tools for Together AI.

FAQs

Do I need a separate Together AI account to use Together through Orq.ai?

Yes. You sign up with Together AI, obtain an API key, then configure it in Orq.ai’s AI Router; Orq.ai then gives you one place to manage routing, observability, and cost controls around that Together usage.

Can I route only some workflows to Together AI and others to different providers?

Yes. You define routes per workflow in Orq.ai and choose which ones should use Together vs other providers, so you can reserve Together for open‑model, cost‑sensitive, or long‑context workloads while sending other tasks to proprietary or regional providers.

Does using Together AI through Orq.ai add latency?

Orq.ai is designed as a lightweight router, so the added overhead is small compared to model inference time. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets while gaining visibility and control.

Alternatives to

Together AI

Anthropic

Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.

Chat

Code

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI

Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Image Generation

Reasoning

Speech

Vision

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI

Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Image Generation

Reasoning

Speech

Vision

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS

Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Reasoning

Vision

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.