
Perplexity
on Orq.ai
Use Perplexity’s Sonar models through a single Orq.ai API. Route models such as Sonar, Sonar Pro, Sonar Reasoning Pro, and Sonar Deep Research via Orq’s AI Router for chat, reasoning, coding, and web‑augmented “ask the internet” workloads.
Capabilities:
Chat
Vision
Models Supported:
sonar
sonar-deep-research
sonar-pro
sonar-reasoning-pro
Provider HQ:
Perplexity AI, based in the United States with a global API platform.
Access Perplexity through Orq.ai’s AI Router
Perplexity is a search‑native AI provider offering the Sonar family of grounded LLMs, which combine web search, citations, and long‑context reasoning behind a managed API. These models are designed for research, complex Q&A, code‑aware answers, and deep, sourced analysis.
Perplexity models available on Orq.ai
Model | Type | Context / scope | Best for | Pricing tier |
|---|---|---|---|---|
Sonar Pro | Chat / reasoning / search‑grounded | Around 200K tokens context | Complex queries, competitive analysis, detailed research where you want deeper search and more citations | Premium – Perplexity pricing shows Sonar Pro at 3 USD / 1M input tokens and 15 USD / 1M output tokens, plus per‑request base fees (roughly 14–34 USD / 1K requests depending on low/medium/high context) |
Sonar Reasoning Pro | Chat / chain‑of‑thought reasoning / search | Around 128K–200K tokens (verify in docs) | Long‑context RAG, multi‑step reasoning, and production workloads that emphasise step‑by‑step analysis with web grounding | Mid‑tier – Sonar Reasoning Pro is listed around 2 USD / 1M input tokens and 8 USD / 1M output tokens, with request fees similar to Sonar Pro for different context sizes |
Sonar / Sonar Small / standard | Chat / fast / cost‑efficient search | Up to roughly 127K tokens | Fast, lower‑cost tasks such as everyday Q&A, simple research, classification with web lookups, and high‑volume support workflows | Cost‑efficient – Sonar (Small) is priced at 1 USD / 1M input tokens and 1 USD / 1M output tokens, with Search API or Sonar request fees around 5–12 USD per 1K requests depending on search context |
Why use Perplexity through Orq.ai?
Capability | Provider / models | Direct | Through Orq.ai |
|---|---|---|---|
Chat & search | Sonar, Sonar Pro, Sonar Reasoning Pro, Sonar Deep Research | Call Perplexity’s Sonar APIs directly for web‑augmented chat, Q&A, and research with citations | Use Perplexity models through Orq.ai’s router, adding central routing, tracing, evals, budgets, and governance around each search‑grounded request |
Code‑aware Q&A | Sonar / Sonar Pro | Use Sonar directly to answer coding questions, explain snippets, and search for code patterns | Route coding+research workloads through Orq.ai, compare Perplexity against non‑search models, and monitor cost, latency, and quality from one control layer |
Long‑context research | Sonar Pro, Sonar Reasoning Pro, Sonar Deep Research | Use Perplexity’s APIs directly for extended multi‑document analysis and deep research sessions | Use Orq.ai to send only workflows that truly need Perplexity’s deep search stack to Sonar, while routing simpler tasks to cheaper models, all under unified observability |
This gives teams a practical way to use Perplexity where it performs best, web‑augmented reasoning and research, while centralising routing, observability, evals, and cost controls across the wider model stack.
Pricing
Plans and API access
Perplexity offers:
the Perplexity app / Pro / Enterprise plans for end‑users, and
the Perplexity API Platform with Sonar, Search, Agent, and Embeddings APIs
For Sonar‑family API usage (2026 public data):
Token pricing (grounded LLMs)
Sonar:
Input tokens: 1 USD / 1M
Output tokens: 1 USD / 1M
Sonar Pro:
Input tokens: 3 USD / 1M
Output tokens: 15 USD / 1M
Sonar Reasoning Pro:
Input tokens: 2 USD / 1M
Output tokens: 8 USD / 1M
Sonar Deep Research:
Input tokens: 2 USD / 1M
Output tokens: 8 USD / 1M
Citation tokens: 2 USD / 1M
Reasoning tokens: 3 USD / 1M
Search queries: 5 USD / 1K requests
Request‑based search fees
For Sonar, Sonar Pro, and Sonar Reasoning Pro, each query also incurs a per‑request fee that depends on search_context_size (low / medium / high):
Sonar: 5 / 8 / 12 USD per 1K requests (low/medium/high).
Sonar Pro: base fees roughly 14 / 24 / 34 USD per 1K requests reported in independent docs
Sonar Reasoning Pro: similar tiered request pricing to Sonar Pro, with higher cost for high context
The Search API (raw search results) is priced separately at 5 USD per 1K requests with no token charges, useful if you only need raw web data and do your own LLM processing.
Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑model rates, quotas, and plan details when using Perplexity via Orq.ai.
Compatible frameworks and tools
Orq.ai exposes Perplexity through:
a Perplexity provider configuration in AI Router / Model Garden, where you paste your Perplexity API key, and
HTTP / OpenAI‑compatible routing, depending on how Sonar is wired into your workflows
That means:
Existing backend workflows that already call Orq’s router (OpenAI‑compatible or HTTP tools) can be updated to route specific tasks to Sonar models without a separate direct integration.
Agents, code assistants, and tools that integrate with Orq.ai can be configured so that search‑augmented or research steps are served by Perplexity, while other steps use different providers, all sharing the same observability and governance layer.
Check the Orq.ai integration docs for the latest supported frameworks and tools for Perplexity.
FAQs
Do I need a separate Perplexity account to use Perplexity through Orq.ai?
Yes. You create a Perplexity API key in the Perplexity console, then configure it in Orq.ai; in some cases, Perplexity usage billed via Orq.ai may also be available depending on your Orq plan and region. In both setups, Orq.ai provides one place to manage routing, observability, and cost controls around that Perplexity usage.
Can I route only some workflows to Perplexity and others to different providers?
Yes. You define routes per workflow in Orq.ai and choose which ones should use Sonar vs other models, so you can reserve Perplexity for research‑heavy, web‑augmented workloads while sending simpler tasks to cheaper or local providers.
Does using Perplexity through Orq.ai add latency?
Orq.ai is designed as a lightweight router, so the added overhead is small compared to the cost of web search and LLM reasoning. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets while gaining visibility and control.
Alternatives to
Perplexity
Anthropic
Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.
Chat
Code
Reasoning
Vision
Models:
claude-opus-5
claude-sonnet-5
claude-fable-5
Open AI
Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Image Generation
Reasoning
Speech
Vision
Models:
gpt-5.5 (EU)
gpt-5.6-luna (EU)
gpt-5.6-sol (EU)
Google AI
Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Image Generation
Reasoning
Speech
Vision
Models:
gemini-3.5-flash-lite (Gemini API)
gemini-3.6-flash (Gemini API)
Gemini 3 Pro Image
AWS
Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.
Chat
Code
Embeddings
Reasoning
Vision
Models:
eu.anthropic.claude-opus-5
global.anthropic.claude-opus-5
us.anthropic.claude-opus-5


