Overall result quality versus cost across AI search providers (Exa, Keenable, Firecrawl, etc)
This report positions seven AI/web search providers on 13 benchmark observations (recall %, LLM-judged relevance scores on 0–100 scales, 20–50 question samples) against official and community-reported pricing per 1,000 queries. Benchmarks are heterogeneous — the vertical axis mixes metrics and is not a league table.
Includes up to 10 results with token-efficient page contents. Deep Search $12–15/1k; Contents $1/1k pages; Answer $5/1k. Pay-as-you-go, official schedule.
exa.aiCommunity benchmark: cheapest API in test, 696 context tokens per correct answer, 100% recall with snippet+description output.
krabarena.com$50/month for 50,000 Google searches. No JavaScript rendering; relevance depends on Google's ranked results.
aitoolsatlas.ai$5/1k requests search fee plus $1/M input and $5/M output tokens (Sonar tier). Synthesised answers with citations; total cost less predictable.
bestwebsearchapis.comOperational estimate: ~$0.05–0.08 for a 13-test SanityWebEval default run. Native search + extract + crawl — the page-extraction specialist, not a like-for-like semantic engine.
github.comOfficial plans list $5 per 1,000 requests and $4 per 1,000 requests, plus $5 per million input/output tokens for AI-oriented use.
brave.comFree tier 1,000 API credits/month; Project plan 4,000 credits/month; Enterprise custom. Balanced agent-search API.
tavily.comBasic outscored Turbo 73 vs 67 while Turbo ran ~2× faster; captured comparison pages cite call pricing from $1 to $2,400 per 1,000 depending on tier and task depth.
the-decoder.com| Provider | Test | Result | Source |
|---|---|---|---|
| Parallel (turbo) | New benchmark ranks search APIs for AI agents on quality, cost, and speed | 67 | the-decoder.com |
| Exa | Tavily beats Parallel by 2.18 points on 50 random GPT Mini-judged FreshQA retrieval | 0 | krabarena.com |
| Keenable | Tavily beats Parallel by 2.18 points on 50 random GPT Mini-judged FreshQA retrieval | 63.7 | krabarena.com |
| Parallel | Tavily beats Parallel by 2.18 points on 50 random GPT Mini-judged FreshQA retrieval | 69.3 | krabarena.com |
| Tavily | Tavily beats Parallel by 2.18 points on 50 random GPT Mini-judged FreshQA retrieval | 71.48 | krabarena.com |
| Exa | Tavily beats Parallel by 2.35 points on GPT Mini-judged FreshQA retrieval | 59.95 | krabarena.com |
| Keenable | Tavily beats Parallel by 2.35 points on GPT Mini-judged FreshQA retrieval | 56.23 | krabarena.com |
| Parallel | Tavily beats Parallel by 2.35 points on GPT Mini-judged FreshQA retrieval | 61.75 | krabarena.com |
| Tavily | Tavily beats Parallel by 2.35 points on GPT Mini-judged FreshQA retrieval | 64.1 | krabarena.com |
| Exa | Tested 3 search APIs on 50 SimpleQA questions — Exa wins recall (64%), Keenable is 10× faster | 64% | krabarena.com |
| Keenable | Tested 3 search APIs on 50 SimpleQA questions — Exa wins recall (64%), Keenable is 10× faster | 52% | krabarena.com |
| Exa | Tested 4 search APIs on 20 HN-style SimpleQA questions — You.com wins recall (90%), Keenable is 2× faster | 85% | krabarena.com |
| Keenable | Tested 4 search APIs on 20 HN-style SimpleQA questions — You.com wins recall (90%), Keenable is 2× faster | 65% | krabarena.com |
| Tavily | Tested 4 search APIs on 20 HN-style SimpleQA questions — You.com wins recall (90%), Keenable is 2× faster | 80% | krabarena.com |
| Exa | coding and general QA evaluations; complete agent trajectories; BrowseComp, WideSearch, and internal company a | about 49% better token efficiency and 2.4% higher downstream quality with Exa Auto; about 30% fewer tokens across comple | exa.ai |
| Keenable | context tokens per correct answer | 696 tokens | krabarena.com |
Compiled from 22 provider/benchmark observations plus fetched official pricing pages and extracted claim snippets, captured as of 2026-09-01. Results mix retrieval recall, LLM-judged relevance (0–100), downstream token counts, and product tiers over 20–50 question samples; vendor claims and community benchmarks are labelled and no composite score is computed. Non-scoring capability notes were cut for space.