Keenable reports the largest comparable document count among providers that disclose one
A scan of AI-oriented search providers found three current first-party numerical disclosures suitable for a directional comparison. Keenable’s 100B+ documents is 25% above Exa’s 80B documents served and at least 2.5× Brave Search’s 40B+ webpages, but the counting definitions are not identical.
100B+
Keenable documents
1.25×
Keenable / Exa reported count
2.5×
Keenable / Brave minimum
Blue intensity indicates rank
0255075100B
What the disclosures support
1
Keenable leads on the headline count. Its site reports 100B+ documents, the largest directly comparable first-party figure found. keenable.ai
2
Exa is closest at 80B served documents. That puts Keenable’s stated floor 20B higher, or 1.25× Exa’s figure. Exa separately says its crawlers track 500B+ URLs, but tracked URLs are not the same as served documents. exa.ai
3
Brave reports 40B+ webpages. Keenable’s stated document count is therefore at least 2.5× Brave’s stated page count. brave.com
4
Absence is not zero. Other AI search services were not ranked because a current, defensible first-party count for their own index was not found; some depend partly or wholly on upstream indexes.
Reported figures side by side
| Provider | Reported scale | What is counted | Relative to Keenable | Source |
|---|---|---|---|---|
| Keenable | 100B+ documents | Documents in index | Baseline | keenable.ai |
| Exa | 80B documents served | Documents served in vector database | Keenable ≥1.25× | exa.ai |
| Brave Search | 40B+ webpages | Webpages in independent index | Keenable ≥2.5× | brave.com |
Caution: “documents,” “documents served,” “webpages,” and “URLs tracked” can differ through deduplication, canonicalization, versioning, exclusions, and storage tiers. These are provider-reported scale claims, not independently audited apples-to-apples measurements.
Method: searched provider sites, documentation, company blogs, and reporting across more than 1,700 unique search results on 26 August 2026; retained current first-party numerical disclosures for a provider’s own index or serving corpus. Bar lengths use the stated lower-bound number, not an estimate of the unknown amount above “+”.