Ten open datasets cover the field: FineWeb leads English scale at 18.5T+ tokens, and the right pick depends on the job

Asked:
“Which large, freely available open datasets are best suited for fine-tuning or retrieval-augmented generation with open-weight language models, and what makes each one valuable? Include publisher, size, and license.”

Ten distinct datasets, drawn from 13 sourced rows reviewed on 2026-10-08 — official dataset cards on Hugging Face and NCBI pages. Four rows describe the same PMC family through different official pages and are consolidated here into one canonical PMC Open Access Subset entry. Sizes stay in each dataset’s native unit — tokens, documents, files, samples, articles, disk — because they are not directly comparable.

Decision matrix: the shortlist by best fit

Permissive (Apache 2.0) Attribution license, but upstream source terms still apply License varies by subset, file, or article — filter before use Hover or tap a card for publisher, exact size and why it matters
Before commercial training or indexing, filter by per-record or per-subset license. “Open” and “freely downloadable” do not imply uniform downstream rights: FineWeb, FineWeb2 and Dolma are ODC-By yet remain subject to source terms; RedPajama-V2 is governed by Common Crawl terms; CulturaX follows mC4 and OSCAR terms; The Pile varies by subset; The Stack v2 requires compliance with original repository licenses; Tulu 3 contains subset-specific and some non-commercial restrictions; PMC licenses vary article by article. Only Aya Collection is Apache 2.0 throughout, stating its components were selected for permissive manipulation and redistribution.

What the matrix cannot say alone

RedPajama-V2 is the only corpus shipping quality signals at raw scale: 30B of its 100B+ documents carry annotations, plus IDs for a 20B-document deduplicated set — you filter to your own recipe. huggingface.co

FineWeb2 reaches over 1,000 languages (1,868 language-script pairs) with per-language deduplication and tuned filtering — broader coverage than CulturaX’s 167 languages, which in turn offers more tokens (6.3T). huggingface.co

Aya Collection’s 513,758,189 instances dwarf Tulu 3’s 939,344 samples, but Tulu 3 is the curated recipe — 18 sources spanning chat, math, code, instruction following and safety — when you want compact, proven SFT. huggingface.co

For biomedical RAG, the PMC Open Access Subset serves millions of full-text articles in XML or plain text via bulk packages and APIs, with licenses grouped into commercial-use-allowed, non-commercial-only, and other. ncbi.nlm.nih.gov

Every dataset in detail

DatasetBest fitPublisherReported sizeLicenseWhy it mattersSource

Method: official dataset cards and NCBI pages reviewed live on 2026-10-08; 13 sourced rows across two independent hosts (huggingface.co and ncbi.nlm.nih.gov), no URL supporting more than one row. Four PMC-family rows were consolidated into one canonical PMC Open Access Subset entry, giving 10 distinct datasets. Sizes are as reported by publishers in native units (tokens, documents, files, samples, articles, disk size) and are not interconverted. License summaries are abbreviated for space; the full statements are on each linked page.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT