orq.ai Engineering · October 5, 2026
Ask a coding agent what changed in the latest Go release. It needs the current release notes to answer. Training data may be out of date, and the notes may be missing from the documents you’ve supplied.
A web search API lets the agent look them up during the request. It sends a query and gets links and snippets it can use in its answer, along with sources the user can check. Without that lookup, it may invent a URL or leave you to find and paste the page yourself.
Search is useful when an answer depends on public information you don’t already have, such as recent news or developer documentation. If your existing context answers the question, searching again can just add cost and latency.

Adding search providers means managing API keys, retries, spend tracking and PII handling. Web search is now part of the orq.ai AI Gateway: one endpoint gives you Exa, Ceramic, Linkup, Tavily and Serper, using the same credits as your model calls.
Searches appear alongside model calls in traces and use the same PII redaction and guardrails as your prompts. To help you choose a provider, we ran 768 searches to compare relevance, latency and cost.
How it works
Send a query, provider and result limit. Every provider returns url, title and description fields, plus metadata identifying the provider:
Managed searches use your gateway credits at the provider’s list price plus $0.001 per search. With your own provider key, orq.ai does not bill the search.
Each web_search span records the provider, result count, latency and cost, linking the search to the answer it fed. Volume, failures, latency and spend are available by workspace and project, like model metrics.
The PII Redaction plugin masks personal data before the query leaves orq.ai. Input and output guardrails can block queries or results. The endpoint and web_search server tool share a Go provider library that other orq.ai runtimes can reuse; providers only need to be added once.
The benchmark
We tested 48 queries across general knowledge, current news, developer documentation, research papers, local and shopping, and non-English searches (Spanish, Dutch, German, French, Italian and Japanese). Each provider received every query with a 10-result limit and our integration’s defaults: Exa auto with highlights, Linkup standard, Tavily basic, Serper Google web results and Ceramic’s default search.
We pooled, deduplicated and shuffled each query’s URLs, hiding provider names. Claude Opus 5.5 graded each URL, title and snippet from 0 (irrelevant) to 3 (ideal destination). nDCG@10 rewards relevant results near the top; precision@5 counts top-five results graded 2 or higher. A second grading of the same 1,777 URLs agreed 90% of the time; no grade changed by more than one point.
Searches | Queries | URLs graded blind | Provider spend (528 direct searches) | Median gateway overhead |
|---|---|---|---|---|
768 | 48 in 6 categories | 1,947 | $2.05 | 3.1 ms |
Relevance and price
Exa led on relevance, with a good top-five result for every query. Serper came close to Tavily at one eighth of its list price. Linkup placed fourth. Ceramic cost 28 times less than Exa but scored far lower on natural-language queries.

48 queries, round one, pooled and graded blind. Logarithmic price axis. The hollow marker shows Ceramic with the keyword rewrites its documentation recommends.
Provider | nDCG@10 | Precision@5 | Queries with no relevant top-5 result |
|---|---|---|---|
Exa | 0.755 | 93% | 0 of 48 |
Serper | 0.625 | 84% | 1 of 48 |
Tavily | 0.619 | 84% | 0 of 48 |
Linkup | 0.511 | 69% | 2 of 48 |
Ceramic | 0.245 | 30% | 20 of 47 |
Ceramic + keyword rewrite | 0.264 | 34% | 17 of 47 |

Latency
Ceramic’s median was 164 ms; its slowest request took 506 ms. The other providers’ cold-query medians ranged from 1.20 to 2.65 seconds. Linkup had the tightest spread among those four. Exa had the longest tail: auto decides how much work each query needs, with some taking over four seconds.

Client-side latency for first-time queries. Four workers called providers in random order. The chart shows the median, p90 and maximum for each provider.
Provider | p50 | p90 | Max |
|---|---|---|---|
Ceramic | 164 ms | 215 ms | 506 ms |
Linkup | 1.20 s | 1.57 s | 2.06 s |
Serper | 1.33 s | 3.08 s | 5.81 s |
Tavily | 1.56 s | 3.46 s | 3.63 s |
Exa | 2.65 s | 4.67 s | 6.32 s |
Repeats hit upstream caches
Repeating all 48 queries sped up three providers. Tavily’s median fell from 1.56 s to 115 ms; Exa’s fell from 2.65 s to 218 ms. Agents benefit on retries and repeated searches. Use fresh queries for benchmarks, or you’ll mostly measure caches.

Local gateway overhead: 3.1 ms
A third pass sent all 48 queries to every provider through /v3/websearch. Client time minus provider wait time gave a median overhead of 3.1 ms and p90 of 3.9 ms for authentication, validation, billing, metrics and tracing. These measurements used a local gateway and exclude the network hop to orq.ai.
Choosing a provider
Exa led in five of six categories, especially current news. Tavily and Serper tied on non-English queries, with Linkup close behind. Tavily placed second on local and shopping; Linkup trailed on news and research.
Ceramic’s strongest categories were research and general knowledge, where queries tend to use precise terms. Its English keyword engine suits frequent narrow lookups, including parallel searches. News and non-English queries were its weakest categories; it rejected a Japanese query.

Coverage depends on the engine
Most provider pairs shared only 2% to 12% of URLs. Tavily and Serper overlapped by 70% on average, so pairing them adds little coverage. Combining engines with low overlap, such as Exa and Serper, returns more distinct material than switching between them.

Snippets affect the next step
Exa’s highlights had a median of 3,173 characters per result, often enough to answer without fetching the page. Linkup also returned long snippets. Serper’s roughly 146-character Google snippets make it cheap for discovery, but agents that need page content must fetch it separately. Choose Exa when relevance matters more than latency; Serper when cost matters more.

Pricing
The table includes the $0.001 managed surcharge and the cost of 1,000 relevant top-five results. Exa’s reported cost matched its list price exactly.
Provider | Mode | List price / 1k | Measured / 1k | orq.ai managed / 1k | Bring your own key | nDCG@10 | Per 1k relevant top-5 results |
|---|---|---|---|---|---|---|---|
Ceramic | search | $0.25 | $0.25 | $1.25 | $0 on orq.ai | 0.245 | $0.84 |
Serper | google web | $1.00 | $1.00 | $2.00 | $0 on orq.ai | 0.625 | $0.48 |
Linkup | standard | $5.00 | $5.00 | $6.00 | $0 on orq.ai | 0.511 | $1.75 |
Exa | auto + highlights | $7.00 | $7.00 | $8.00 | $0 on orq.ai | 0.755 | $1.72 |
Tavily | basic | $8.00 | $8.00 | $9.00 | $0 on orq.ai | 0.619 | $2.14 |
These are published pay-as-you-go prices on October 5, 2026. Volume plans can cost less: Ceramic Pro is $0.20 per 1,000 searches. The web_search server tool in chat completions and responses costs a flat $0.005 per call.
Limits of the benchmark
48 queries separate broad differences, but cannot reliably distinguish close scores such as Tavily and Serper’s.
The grader saw URLs, titles and snippets, not full pages. Longer snippets may slightly favor Exa and Linkup.
We tested defaults. Exa
deep, Linkupdeepand Tavilyadvancedcost more and may score higher.Latency comes from one location on one afternoon. Your network path will differ.
Try it
Every workspace with AI Gateway credits can use web search. Choose a provider and send your orq.ai API key:
To use your own key, add the provider account as an integration. Searches still appear in traces and metrics, without an orq.ai search charge.




