As of 2026-09-01

No universal winner: Exa leads on recall, Keenable on cost per useful token, Tavily on judged relevance

Asked:

Overall result quality versus cost across AI search providers (Exa, Keenable, Firecrawl, etc)

This report positions seven AI/web search providers on 13 benchmark observations (recall %, LLM-judged relevance scores on 0–100 scales, 20–50 question samples) against official and community-reported pricing per 1,000 queries. Benchmarks are heterogeneous — the vertical axis mixes metrics and is not a league table.

Quality signals vs cost per 1,000 queries

Retrieval recall (% correct, 20–50 q samples) LLM-judged relevance (0–100, GPT Mini judge) Hollow ring = cost not captured in the same test; plotted in the unpriced band

What the scatter cannot say

Exa's own changelog reports Auto delivering about 49% better token efficiency and 2.4% higher single-turn downstream quality, and about 30% fewer tokens with 1% higher quality across full agent trajectories — vendor-reported figures, not third-party. exa.ai
A community token-tax benchmark found Keenable the most efficient real-content API: 696 context tokens per correct answer at 100% recall for $1/1k, while Exa's full-page mode billed an agent 43× more tokens per search. krabarena.com
The same benchmark's leanest result — 217 tokens per question at 100% recall — used SERP descriptions only, so it may degrade on questions that need page-body content; Firecrawl and Keenable were the joint efficiency winners. krabarena.com
Parallel Turbo scored 67 vs 73 for Parallel Basic in a quality/cost/speed benchmark while running roughly twice as fast (0.51s vs 1.03s per the source) — Turbo trades quality for speed and can require more passes. the-decoder.com
Serper's $50/month for 50,000 Google searches ($1/1k) is the cheapest SERP layer, but it does not render JavaScript and quality depends on Google's ranking. aitoolsatlas.ai

Pricing, from official pages and captured estimates

Exa

$7 / 1k searches

Includes up to 10 results with token-efficient page contents. Deep Search $12–15/1k; Contents $1/1k pages; Answer $5/1k. Pay-as-you-go, official schedule.

exa.ai

Keenable

$1 / 1k

Community benchmark: cheapest API in test, 696 context tokens per correct answer, 100% recall with snippet+description output.

krabarena.com

Serper

$1 / 1k

$50/month for 50,000 Google searches. No JavaScript rendering; relevance depends on Google's ranked results.

aitoolsatlas.ai

Perplexity Sonar

$5 / 1k + tokens

$5/1k requests search fee plus $1/M input and $5/M output tokens (Sonar tier). Synthesised answers with citations; total cost less predictable.

bestwebsearchapis.com

Firecrawl

$0.005 / credit

Operational estimate: ~$0.05–0.08 for a 13-test SanityWebEval default run. Native search + extract + crawl — the page-extraction specialist, not a like-for-like semantic engine.

github.com

Brave

$4–5 / 1k

Official plans list $5 per 1,000 requests and $4 per 1,000 requests, plus $5 per million input/output tokens for AI-oriented use.

brave.com

Tavily

Credit plans

Free tier 1,000 API credits/month; Project plan 4,000 credits/month; Enterprise custom. Balanced agent-search API.

tavily.com

Parallel

Tiered

Basic outscored Turbo 73 vs 67 while Turbo ran ~2× faster; captured comparison pages cite call pricing from $1 to $2,400 per 1,000 depending on tier and task depth.

the-decoder.com

Which one to choose

ExaQuality-first semantic search and deep research; budget for the downstream token tax of full-page output.
KeenableCompact structured multi-source research with the lowest downstream context burden.
TavilyA balanced agent-search API — strong judged relevance without a premium price posture.
FirecrawlWhen reliable page extraction and crawling is the dominant requirement.
ParallelSpeed and deep-task workflows; pick the tier deliberately — Turbo trades quality for speed.
SerperThe cheapest raw Google SERP feed when links and snippets are enough.
Perplexity SonarAn answer API with citations, if variable token billing is acceptable.
Before committingRun a 100–300 query bake-off on your own query distribution; total cost = API fee + retries + returned-context tokens + LLM tokens + rendering needs.

Every benchmark row

ProviderTestResultSource
Parallel (turbo)New benchmark ranks search APIs for AI agents on quality, cost, and speed67the-decoder.com
ExaTavily beats Parallel by 2.18 points on 50 random GPT Mini-judged FreshQA retrieval0krabarena.com
KeenableTavily beats Parallel by 2.18 points on 50 random GPT Mini-judged FreshQA retrieval63.7krabarena.com
ParallelTavily beats Parallel by 2.18 points on 50 random GPT Mini-judged FreshQA retrieval69.3krabarena.com
TavilyTavily beats Parallel by 2.18 points on 50 random GPT Mini-judged FreshQA retrieval71.48krabarena.com
ExaTavily beats Parallel by 2.35 points on GPT Mini-judged FreshQA retrieval59.95krabarena.com
KeenableTavily beats Parallel by 2.35 points on GPT Mini-judged FreshQA retrieval56.23krabarena.com
ParallelTavily beats Parallel by 2.35 points on GPT Mini-judged FreshQA retrieval61.75krabarena.com
TavilyTavily beats Parallel by 2.35 points on GPT Mini-judged FreshQA retrieval64.1krabarena.com
ExaTested 3 search APIs on 50 SimpleQA questions — Exa wins recall (64%), Keenable is 10× faster64%krabarena.com
KeenableTested 3 search APIs on 50 SimpleQA questions — Exa wins recall (64%), Keenable is 10× faster52%krabarena.com
ExaTested 4 search APIs on 20 HN-style SimpleQA questions — You.com wins recall (90%), Keenable is 2× faster85%krabarena.com
KeenableTested 4 search APIs on 20 HN-style SimpleQA questions — You.com wins recall (90%), Keenable is 2× faster65%krabarena.com
TavilyTested 4 search APIs on 20 HN-style SimpleQA questions — You.com wins recall (90%), Keenable is 2× faster80%krabarena.com
Exacoding and general QA evaluations; complete agent trajectories; BrowseComp, WideSearch, and internal company aabout 49% better token efficiency and 2.4% higher downstream quality with Exa Auto; about 30% fewer tokens across compleexa.ai
Keenablecontext tokens per correct answer696 tokenskrabarena.com

Compiled from 22 provider/benchmark observations plus fetched official pricing pages and extracted claim snippets, captured as of 2026-09-01. Results mix retrieval recall, LLM-judged relevance (0–100), downstream token counts, and product tiers over 20–50 question samples; vendor claims and community benchmarks are labelled and no composite score is computed. Non-scoring capability notes were cut for space.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Made withKeenable SELECT · 4,654 pages in 6m 43s · Ask your own questionShare:XLinkedInReddit
Made with Keenable SELECT