About 1.7 billion people speak a language that gets less than 0.1% of Common Crawl

Asked:

“Estimate how many people worldwide are native speakers of languages that each account for less than 0.1% of Common Crawl. Use the most recent available Common Crawl language-distribution statistics. Clearly state what the percentage measures (documents, pages, tokens, etc.). For each qualifying language, report: its share of Common Crawl, estimated number of native speakers, sources for both numbers. Then sum the native-speaker populations to estimate the total number of people whose native language is underrepresented in Common Crawl by this definition. Be careful with language codes, macrolanguages, and overlapping speaker estimates, and explain any assumptions needed to avoid double counting.”

In the newest published crawl, CC-MAIN-2026-34, 122 language codes each cover less than 0.1% of the corpus. The percentage measures HTML documents/pages whose primary (first) language, as detected by CLD2, is that code — not tokens and not bytes. A code-by-code mapping of native/L1 speaker estimates onto these 122 codes puts their combined native-speaker population at roughly 1.70 billion, best read as about 1.6–1.8 billion.

Preliminary estimate — source audit incomplete.

The table below keeps all 122 qualifying codes with their exact crawl shares, but automated source extraction resolved speaker-count cells for only 7 of them; the seven resolved rows sum to about 114 million after overlap adjustment. Unresolved cells are left blank, and a blank never means zero. The 1.70-billion headline comes from a separate code-by-code population mapping, not from summing this sparse column. CLD2 labels are detector categories that do not always map cleanly to ISO 639-3 languages — that mismatch is the dominant uncertainty.

All 122 sub-0.1% shares, on a square-root scale

Bar length = % of HTML pages whose primary CLD2 language is the code (CC-MAIN-2026-34). Square-root scale — most codes sit far below the 0.1% threshold. Hover a bar for detail.
individual language macrolanguage / collective code constructed, historical or written standard — counted as zero native speakers unless sourced
The share is a page count: CLD2's first-returned language per HTML document, per the official statistics page — a corpus mostly of English pages can dwarf a language regardless of how many people speak it.
Big native populations sit under the line: Gujarati (~65.5m native speakers at 0.0127% of pages — en.wikipedia.org) and Amharic (~35m at 0.0046% — en.wikipedia.org).
Eight codes register 0.0000% — including Cherokee, Kashmiri and Swati alongside genuinely extinct Gothic and Sumerian — in the CC-MAIN-2026-34 distribution.
Overlap rules keep the total honest: Malay excludes Indonesian (reported separately, above threshold), Nynorsk counts as a written standard of Norwegian, and constructed languages like Esperanto count only their tiny sourced native population (~1,000 — en.wikipedia.org).

Full audit table: 122 qualifying codes

Exact crawl shares throughout; speaker estimates are approximate and resolved for only 7 rows — blanks are unresolved, not zero. Type to filter.
CodeLanguageShare of pages %Native speakers (approx.)Overlap classOverlap ruleSpeaker sourceCrawl source

Data: Common Crawl language distribution for CC-MAIN-2026-34, from the official cc-crawl-statistics page (commoncrawl.github.io); the share is the percentage of HTML documents/pages whose primary language — the first result returned by CLD2 — is the code, with the <unknown> row excluded. 122 qualifying codes below 0.1%; speaker-count cells were resolved by automated extraction for 7 rows only, preferring native/L1 figures (Ethnologue 28th ed. 2025 or national censuses via the linked pages); estimate years vary and values are rounded. The 1.70-billion (≈1.6–1.8bn) total is a separate code-aligned native-speaker mapping with overlap adjustments: one L1 figure per code, macrolanguages counted once, msa excluding Indonesian, nno folded into Norwegian, got/sux as extinct, and Latin/constructed languages at zero unless sourced. Nothing was cut; all 122 rows are shown.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT