Jina

on Orq.ai

Use Jina’s search and retrieval APIs through a single Orq.ai endpoint. Route Jina’s embeddings, rerankers, and Reader API via Orq’s AI Router to power RAG, semantic search, and retrieval‑heavy workloads alongside your LLM stack.

Capabilities:

Chat

Embeddings

Speech

Models Supported:

jina-embeddings-v5-omni-nano

jina-embeddings-v5-omni-small

jina-embeddings-v5-text-nano

jina-embeddings-v5-text-small

jina-reranker-v3

Provider HQ:

Jina AI headquartered in Sunnyvale, California, USA.

Access Jina through Orq.ai’s AI Router

Jina AI provides a suite of “search foundation” APIs: multimodal embeddings, late‑interaction rerankers, and a Reader API that turns web pages or documents into clean text/Markdown, all with token‑based billing shared across products. These services are designed to be a complete retrieval layer for production RAG pipelines at low per‑token cost.

Jina models available on Orq.ai

Why use Jina through Orq.ai?

Capability

Provider / models

Direct

Through Orq.ai

Embeddings

Jina embedding APIs (text / multimodal)

Call Jina’s embedding endpoints directly to generate vectors for text and multimodal content

Use Jina embeddings through Orq.ai’s router so retrieval stays observable and configurable alongside your LLM requests

Rerank

Jina Reranker V2 Base (multilingual) and related

Use Jina rerankers directly to reorder search/RAG results based on neural relevance

Route rerank calls via Orq.ai and keep logs, metrics, and evals in the same place as your model traffic

Reader

Reader API (web‑to‑Markdown)

Call Reader directly to turn URLs and documents into clean text/Markdown for indexing

Integrate Reader into Orq‑managed workflows so the same router that handles LLM calls also orchestrates content ingestion

This gives teams a practical way to use Jina where it performs best while centralising routing, observability, evals, and cost controls across the wider model stack.

Pricing

Model rates

Jina pricing is token‑based and shared across APIs (embeddings, rerankers, Reader):

  • Free / non‑commercial tier: around 1–10M tokens included for experimentation.

  • Prototype tier: about 0.05 USD per 1M tokens (good for small‑scale production).

  • Production tier: about 0.045 USD per 1M tokens for high‑volume enterprise RAG.

Specific APIs like Jina Reranker V2 Base are listed at about 0.018 USD per 1M input and 0.018 USD per 1M output tokens.

Check the Orq.ai pricing page and your workspace’s provider configuration for current per‑API rates, quotas, and plan details.

Compatible frameworks and tools

Orq.ai exposes Jina via:

  • a Jina provider configuration in the AI Router / Model Garden, and

  • standard HTTP/OpenAI‑compatible integrations in supported frameworks.

That means:

  • Popular RAG stacks (for example Qdrant, LangChain‑style frameworks, AI SDKs) can call Jina embeddings and rerankers via Orq’s router instead of wiring each integration separately.

  • Agents, code assistants, and tools that already talk to Orq.ai (through OpenAI‑compatible or HTTP tools) can incorporate Jina for retrieval without changing their core integration.

Check the Orq.ai integration docs for the latest supported frameworks and tools for Jina.

FAQs

Do I need a separate Jina account to use Jina through Orq.ai?

You can either connect your own Jina API key into Orq.ai or, where available, use Jina usage billed via Orq.ai; the exact options depend on your Orq plan, region, and how Jina is configured in your workspace. In both cases, Orq.ai gives you one place to manage routing, observability, and cost controls around that Jina usage.

Can I route only some workflows to Jina and others to different providers?

Yes. You define routes per workflow in Orq.ai and decide which ones should use Jina vs other retrieval or embedding providers, so you can reserve Jina for specific RAG paths while sending other tasks to different stacks.

Does using Jina through Orq.ai add latency?

Orq.ai is designed as a lightweight router layer, so the added overhead is small compared to embedding, rerank, or Reader processing times. You can use routing policies, caching, and provider selection to keep end‑to‑end performance within your targets.

Alternatives to

Jina

Anthropic

Use Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, and other supported Claude models through one API.

Chat

Code

Reasoning

Vision

Models:

claude-opus-5

claude-sonnet-5

claude-fable-5

Open AI

Use OpenAI's foundation models through a single Orq.ai API. Route models such as GPT-4.1, GPT-4.1-mini, o3-mini, and GPT-4o-class models via Orq's AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Image Generation

Reasoning

Speech

Vision

Models:

gpt-5.5 (EU)

gpt-5.6-luna (EU)

gpt-5.6-sol (EU)

Google AI

Use Google’s Gemini models through a single Orq.ai API. Route models such as Gemini 3.1 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash‑Lite via Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Image Generation

Reasoning

Speech

Vision

Models:

gemini-3.5-flash-lite (Gemini API)

gemini-3.6-flash (Gemini API)

Gemini 3 Pro Image

AWS

Use AWS Bedrock’s foundation models through a single Orq.ai API. Route models such as Amazon Nova, Amazon Titan Text, and compatible third‑party models exposed via Bedrock through Orq’s AI Router for chat, reasoning, coding, and multimodal workloads.

Chat

Code

Embeddings

Reasoning

Vision

Models:

eu.anthropic.claude-opus-5

global.anthropic.claude-opus-5

us.anthropic.claude-opus-5

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.

Orq.ai AI Gateway gradient artwork

Get your API key and start routing in minutes

€1 of free credit included. No card. Live in two minutes. The full platform is there when you need it.