“Estimate how many people worldwide are native speakers of languages that each account for less than 0.1% of Common Crawl. Use the most recent available Common Crawl language-distribution statistics. Clearly state what the percentage measures (documents, pages, tokens, etc.). For each qualifying language, report: its share of Common Crawl, estimated number of native speakers, sources for both numbers. Then sum the native-speaker populations to estimate the total number of people whose native language is underrepresented in Common Crawl by this definition. Be careful with language codes, macrolanguages, and overlapping speaker estimates, and explain any assumptions needed to avoid double counting.”
In the newest published crawl, CC-MAIN-2026-34, 122 language codes each cover less than 0.1% of the corpus. The percentage measures HTML documents/pages whose primary (first) language, as detected by CLD2, is that code — not tokens and not bytes. A code-by-code mapping of native/L1 speaker estimates onto these 122 codes puts their combined native-speaker population at roughly 1.70 billion, best read as about 1.6–1.8 billion.
The table below keeps all 122 qualifying codes with their exact crawl shares, but automated source extraction resolved speaker-count cells for only 7 of them; the seven resolved rows sum to about 114 million after overlap adjustment. Unresolved cells are left blank, and a blank never means zero. The 1.70-billion headline comes from a separate code-by-code population mapping, not from summing this sparse column. CLD2 labels are detector categories that do not always map cleanly to ISO 639-3 languages — that mismatch is the dominant uncertainty.
| Code | Language | Share of pages % | Native speakers (approx.) | Overlap class | Overlap rule | Speaker source | Crawl source |
|---|
Data: Common Crawl language distribution for CC-MAIN-2026-34, from the official cc-crawl-statistics page (commoncrawl.github.io); the share is the percentage of HTML documents/pages whose primary language — the first result returned by CLD2 — is the code, with the <unknown> row excluded. 122 qualifying codes below 0.1%; speaker-count cells were resolved by automated extraction for 7 rows only, preferring native/L1 figures (Ethnologue 28th ed. 2025 or national censuses via the linked pages); estimate years vary and values are rounded. The 1.70-billion (≈1.6–1.8bn) total is a separate code-aligned native-speaker mapping with overlap adjustments: one L1 figure per code, macrolanguages counted once, msa excluding Indonesian, nno folded into Norwegian, got/sux as extinct, and Latin/constructed languages at zero unless sourced. Nothing was cut; all 122 rows are shown.