Ten distinct datasets, drawn from 13 sourced rows reviewed on 2026-10-08 — official dataset cards on Hugging Face and NCBI pages. Four rows describe the same PMC family through different official pages and are consolidated here into one canonical PMC Open Access Subset entry. Sizes stay in each dataset’s native unit — tokens, documents, files, samples, articles, disk — because they are not directly comparable.
RedPajama-V2 is the only corpus shipping quality signals at raw scale: 30B of its 100B+ documents carry annotations, plus IDs for a 20B-document deduplicated set — you filter to your own recipe. huggingface.co
FineWeb2 reaches over 1,000 languages (1,868 language-script pairs) with per-language deduplication and tuned filtering — broader coverage than CulturaX’s 167 languages, which in turn offers more tokens (6.3T). huggingface.co
Aya Collection’s 513,758,189 instances dwarf Tulu 3’s 939,344 samples, but Tulu 3 is the curated recipe — 18 sources spanning chat, math, code, instruction following and safety — when you want compact, proven SFT. huggingface.co
For biomedical RAG, the PMC Open Access Subset serves millions of full-text articles in XML or plain text via bulk packages and APIs, with licenses grouped into commercial-use-allowed, non-commercial-only, and other. ncbi.nlm.nih.gov
| Dataset | Best fit | Publisher | Reported size | License | Why it matters | Source |
|---|
Method: official dataset cards and NCBI pages reviewed live on 2026-10-08; 13 sourced rows across two independent hosts (huggingface.co and ncbi.nlm.nih.gov), no URL supporting more than one row. Four PMC-family rows were consolidated into one canonical PMC Open Access Subset entry, giving 10 distinct datasets. Sizes are as reported by publishers in native units (tokens, documents, files, samples, articles, disk size) and are not interconverted. License summaries are abbreviated for space; the full statements are on each linked page.