Theme
Datasets and Benchmarks Reference
This page is a one-stop reference for data and evaluation benchmarks in the LLM world: pretraining corpora (building the model), post-training instruction data (teaching the model), and evaluation benchmarks (measuring the model). Each entry includes scale, construction method, use case, a one-line description, and licensing notes. The scale and leaderboard numbers for data and benchmarks are continuously updated — this page's data is current as of August 2025; always refer to official sources for the latest figures.
Three paths through this page
Working on pretraining / continued pretraining → see "I. Pretraining Corpora"; Working on fine-tuning / alignment → see "II. Post-training and Instruction Data"; Working on evaluation / model selection → see "III. Evaluation Benchmarks" and "IV. Chinese and Multilingual Data." Data sourcing and cleaning details are covered in Pretraining: Data and Objectives.
I. Pretraining Corpora
Pretraining corpora are "the model's textbooks": scale and quality directly determine the model's capability ceiling. Internet crawl data (Common Crawl) is the raw material for most corpora and must be cleaned, deduplicated, filtered, and blended before it's usable.
| Corpus | Scale | Construction method | One-line description | Licensing notes |
|---|---|---|---|---|
| Common Crawl | ~3–5 billion webpages per month, PB-scale | Full-internet crawling on a monthly basis since 2008 | A freely available web crawl corpus — the "raw ore" for all large-scale pretraining corpora | Data is publicly downloadable, but page content is copyrighted by the originating websites; respect robots.txt and copyright |
| The Pile | ~825 GiB across 22 subsets | Hand-curated mix of books, papers, code, web pages, etc. by EleutherAI | The first "carefully blended" open pretraining corpus, used by models like GPT-NeoX | Licenses vary by subset; includes copyrighted text (e.g. books), so commercial use requires per-subset evaluation |
| RedPajama-Data-1T | ~1.2 trillion tokens | TOGETHER's recreation of the Llama corpus blend | An open reproduction of Llama's 7-way data blend (67% web / 4.5% code / 4.5% books, etc.) | Permissive and commercially usable; both source code and data are open |
| RefinedWeb | Public version: ~600B tokens | Common Crawl cleaned by TII (the Falcon team) | A pure-web corpus emphasizing "high-quality cleaning + rigorous deduplication," used by the Falcon series | Based on Common Crawl; follow the same copyright requirements as the source |
| CulturaX | ~6.3 trillion tokens, 167 languages | Cleaned and deduplicated merge of mC4 and OSCAR | A multilingual corpus covering 167 languages with strong coverage of low-resource non-English languages | Copyright status for multilingual data is complex; evaluate target-language corpora before commercial use |
1. From Raw Crawl to Training Corpus: the Cleaning Pipeline
Raw Common Crawl is extremely noisy (ads, garbled text, duplicates, low-quality pages) and cannot be used for training directly. The standard pipeline has four steps:
- Collection and parsing: Download WARC snapshots, extract body text (strip HTML tags, navigation, ads), and determine language per URL.
- Quality filtering: Remove low-quality pages using rule-based criteria (length, punctuation density, symbol ratio) and heuristic scores (language model perplexity, classifiers). Also filter for toxicity and privacy.
- Deduplication: MinHash approximate deduplication + exact deduplication to remove duplicated content across pages (repeated data wastes compute and causes the model to "memorize test answers").
- Blending and sampling: Mix corpora according to multilingual/code/math/book ratios, then up-sample or down-sample by perplexity.
Why everyone starts from Common Crawl
Common Crawl is free, massive, and continuously updated — it's the "raw material" of data. But the real difference between corpora lies in the cleaning pipeline: deduplication algorithms, quality scorers, and language blends all vary. "Data is the model" mainly refers to these processing details. New corpora (like CulturaX, RefinedWeb) are mostly "better cleaning" rather than "newer sources."
II. Post-training and Instruction Data
Post-training data is orders of magnitude smaller than pretraining corpora (millions vs. trillions of tokens), but quality and diversity determine how "obedient" the model becomes. Note: a large amount of instruction data is generated by closed commercial models, raising licensing and copyright concerns (see the warning box at the end).
| Dataset | Scale | Source | One-line description | Licensing notes |
|---|---|---|---|---|
| Alpaca | ~52K instruction–response pairs | Stanford, generated using text-davinci-003 | The classic open SFT dataset using self-instruct to expand instructions | Generated by an OpenAI model; officially restricted to academic research use |
| ShareGPT | ~90K conversation turns | Community-shared ChatGPT conversation logs | Real user conversations that served as the training source for models like Vicuna | Generated by ChatGPT and shared by users — license is ambiguous, use with caution commercially |
| UltraChat | ~1.5M conversation turns | Tsinghua University, generated via ChatGPT across 30 topics | Large-scale multi-turn conversations covering instructions, queries, and chit-chat | Generated by a commercial model; primarily for research use |
| OpenAssistant (oasst1) | ~16K conversation trees | LAION, crowdsourced human-written | High-quality purely human-written multi-turn conversations — no "AI-generated" controversy | Open source (Apache 2.0), commercially usable |
| Dolly | ~15K entries | Databricks employees, hand-written | Human-written instruction–response pairs, small but clean — great for testing | Open source and commercially usable |
1. Three Ways to Produce Instruction Data
| Method | Examples | Pros | Cons |
|---|---|---|---|
| Human-written | oasst1, Dolly | High quality, clean copyright, commercially usable | High cost, limited scale |
| Model-generated | Alpaca, UltraChat, ShareGPT | Cheap, can scale up massively | Copyright controversies; student models inherit the teacher model's flaws |
| Human-in-the-loop | Most commercial SFT data | A balance between quality and scale | Complex workflow, still not cheap |
2. Scale Reference Values
| Stage | Typical scale | Notes |
|---|---|---|
| Pretraining | ~20–30B tokens for a 1B model; ~2–4T tokens for a 7B model | Chinchilla recommends "approximately 1:20 ratio for parameters to data tokens"; frontier models generally use 10T+ tokens |
| SFT | 1K–100K high-quality instructions | 1,000 high-quality examples can already make a noticeable difference (LIMA experiment); quality and diversity are the keys |
| Preference alignment (RLHF/DPO) | Tens of thousands to hundreds of thousands of preference pairs | Must cover typical failure modes for helpfulness, honesty, and safety |
Quality over quantity is the iron rule
For post-training data, "1,000 high-quality examples" often delivers more value than "100,000 low-quality filler examples." The scale references above are starting points only; your final call should be based on what works on your actual task. See Fine-tuning and Alignment for more.
Compliance red lines for AI-generated data
Alpaca, ShareGPT, UltraChat, and others are generated by OpenAI/Anthropic commercial models — OpenAI's terms of service explicitly prohibit "using outputs to train competing models," and the copyright status of AI-generated content is legally contested. Use only for research and internal experiments; for commercial products, prioritize human-written or licensed data. See Fine-tuning: SFT and PEFT for discussion on data quality and quantity.
III. Evaluation Benchmarks
Evaluation benchmarks are "model exam papers," classified by the capabilities they assess. Keep three things in mind: leaderboard scores are only comparable when using "the same prompts and implementation"; older benchmarks like MMLU may suffer from data contamination; and gaming a single benchmark ≠ real capability (see Evaluation and Benchmarks and Evaluation in Practice).
1. General Knowledge and Language Understanding
| Benchmark | Scale | One-line description |
|---|---|---|
| GLUE | 9 tasks, ~250K samples | A 2018 text-understanding benchmark, the de facto standard of the BERT era, now largely saturated |
| SuperGLUE | 8 more difficult reasoning tasks | GLUE's successor, released in 2019, also now largely "beaten" by LLMs |
| MMLU | 57 disciplines, ~16K multiple-choice questions | The most classic general-knowledge benchmark, long the #1 leaderboard metric |
| HELM | 42 scenarios × 7 metrics | Stanford's multi-metric evaluation framework emphasizing "transparency" and fair comparisons |
2. Math and Reasoning
| Benchmark | Scale | One-line description |
|---|---|---|
| GSM8K | ~8K grade-school math word problems | Elementary/middle school arithmetic, testing multi-step reasoning; used with CoT |
| MATH | ~12.5K competition-level problems | High school to competition-level math across 5 difficulty levels, with strong discrimination |
| BBH | 23 hard tasks from BIG-Bench | Tests logic, deduction, multi-step reasoning, and other "smart" tasks — small models generally can't solve them |
| ARC | ~7.8K science questions | AI2's science QA (ARC-Challenge is harder), testing common sense and scientific knowledge |
3. Code
| Benchmark | Scale | One-line description |
|---|---|---|
| HumanEval | 164 Python function completion tasks | The de facto benchmark for code generation, evaluated with pass@k |
| MBPP | 974 beginner-level Python tasks | A more basic programming benchmark, often reported alongside HumanEval |
4. Long Context and Conversations
| Benchmark | Scale | One-line description |
|---|---|---|
| LongBench | 21 tasks, ~4,750 questions | Bilingual (Chinese/English) long-context evaluation with average input lengths of 10K+ tokens, measuring "how long can the model actually read" |
| MT-Bench | 80 multi-turn open-ended questions | A conversational quality benchmark scored by LLM-as-a-judge (GPT-4); an important reference for Chatbot Arena |
| TruthfulQA | 817 questions | Specifically designed to test "hallucination" and factual accuracy: whether the model produces plausible-sounding but incorrect answers |
5. Instruction Following and More
| Benchmark | Scale | One-line description |
|---|---|---|
| IFEval | ~500 instructions | Specifically tests "whether the model actually follows instructions" (format/constraint adherence), with good discrimination |
| BIG-Bench | 204 tasks | The largest heterogeneous benchmark; in practice, the LLM community mostly uses the harder subset, BBH |
| SQuAD | ~100K questions | A 2016 extractive reading comprehension benchmark, a classic task of the BERT era |
6. Four Common Pitfalls of Evaluation Benchmarks
- Data contamination: Benchmarks like MMLU have been publicly available for years, so questions may have leaked into training data, inflating scores. Defenses include deduplication, monitoring, and building private test sets.
- Saturation and score gaming: GLUE/SuperGLUE are fully saturated; MMLU is approaching saturation — when discrimination drops, the #1 and #10 ranked models may be in the same tier.
- Inconsistent evaluation methodology: Whether CoT prompting is used, how many few-shot examples are provided, and what temperature is set all significantly affect scores. Cross-model comparisons must use the same methodology.
- Leaderboards ≠ real business needs: Benchmarks measure "exam-taking ability." Real-world performance depends on your own task distribution — always build your own evaluation set.
How to choose benchmarks
- General capability comparison: the trio of MMLU + GSM8K + HumanEval (low cost, reproducible);
- Math and reasoning: GSM8K + MATH + BBH;
- Code: HumanEval + MBPP;
- Long context: LongBench (+ build your own "needle in a haystack" task);
- Conversation and "real-world feel": MT-Bench / Chatbot Arena;
- Chinese: see the next section on C-Eval / CMMLU. For building evaluations in practice, see Evaluation in Practice.
IV. Chinese and Multilingual Data
Chinese evaluation and data is a must-have for Chinese developers, so we list them in a dedicated table:
| Dataset | Scale | Use case | One-line description |
|---|---|---|---|
| C-Eval | ~14K questions | Chinese general knowledge | 52 disciplines, covering middle school to graduate level — the Chinese MMLU |
| CMMLU | ~11.5K questions | Chinese general knowledge | 67 disciplines including humanities and social science sub-categories; complements C-Eval |
| CLUE | 9 task types | Chinese language understanding | The Chinese GLUE, the Chinese standard of the BERT era |
| SuperCLUE | Multi-task suite | Chinese conversational ability | A Chinese conversational benchmark with basic/advanced/expert difficulty tiers |
| BELLE | ~1M Chinese instructions | Chinese SFT | Baidu's open-source self-instruct Chinese instruction data |
| COIG | ~190K Chinese instructions | Chinese SFT | Multi-source Chinese instruction data released by BAAI (Beijing Academy of AI) |
Multilingual best practices
For Chinese LLM evaluation, run at least C-Eval + CMMLU (knowledge) + Chinese code/conversation subsets, and check the model's tokenizer efficiency on Chinese — some Western-prioritized models consume significantly more tokens on Chinese input (see Tokenization and Vocabulary).
V. Licensing and Usage Considerations
Copyright and AI-generated data: compliance red lines
Pretraining corpora: Common Crawl and its derivatives containa large amount of copyrighted material, and several countries have seen lawsuits over training data. Commercial products must evaluate data sources and regional laws carefully. AI-generated instruction data: Alpaca, ShareGPT, UltraChat, and others are generated by OpenAI/Anthropic commercial models — OpenAI's terms prohibit "training competing models with outputs," and the copyright of generated content is legally contested. Use only for research and internal experiments. Evaluation benchmarks: Most benchmarks carry research licenses (academic use only); check each dataset's LICENSE before large-scale commercial evaluation.
A four-question compliance checklist
- Is the data source explicitly licensed for your intended use? 2. Does the service terms of any model used to generate the data prohibit using it for training? 3. What are the legal risks in your target jurisdiction (e.g., the EU AI Act, national copyright rulings)? 4. Has your commercial product completed a data-lineage audit? If all four are satisfied, you're clear to launch.
VI. Data Processing Toolchain
| Tool | Use case | Link |
|---|---|---|
| Hugging Face Datasets library | Standard interface for dataset loading, streaming, splitting, and export | https://huggingface.co/docs/datasets |
| datatrove | Large-scale data cleaning pipeline (the tool that cleaned FineWeb) — deduplication, filtering, and scoring end-to-end | https://github.com/huggingface/datatrove |
| lm-evaluation-harness | Standard framework for running evaluation benchmarks like MMLU, GSM8K, and HumanEval | https://github.com/EleutherAI/lm-evaluation-harness |
| OpenCompass | Chinese and multimodal evaluation suite with leaderboard services | https://github.com/open-compass/opencompass |
VII. Five Steps to Build a Data Pipeline from Scratch
A minimal viable workflow for readers who want to use their own data for pretraining or fine-tuning:
- Define goals and budget: Start with the model scale and training token budget (see the "scale reference values" above), then work backward to figure out how much raw data you need — cleaning and deduplication typically discard 50%–90% of raw data.
- Choose data sources: For pretraining, use open corpora (RedPajama/RefinedWeb) or a custom crawler; for instruction data, prioritize human-written or human-in-the-loop data, avoiding unauthorized AI-generated data (see "V. Licensing").
- Clean and deduplicate: Run quality filtering, MinHash deduplication, and toxicity filtering with datatrove; split multilingual corpora by language.
- Blend and sample: Decide on multilingual/code/math/book ratios based on your task objectives; up-sample "useful" corpora using perplexity.
- Validate quality: Start with a small-scale trial run using 100M tokens, observe loss curves and downstream benchmark (C-Eval/MMLU) changes across data-cleaning steps — every step of the data pipeline should be validated by experiments, not intuition.
Further Reading
- Pretraining: Data and Objectives — Complete methods for data cleaning, deduplication, and corpus blending
- Evaluation and Benchmarks — Benchmark limitations, data contamination, and a critical view of leaderboards
- Evaluation in Practice — Building an evaluation pipeline with lm-eval-harness
- Fine-tuning: SFT and PEFT — How instruction data is used for SFT
- RAG: Retrieval-Augmented Generation — An alternative path for deploying private data (without training)
- Curated Resource List — A full toolchain of datasets and evaluation tools
References
- Common Crawl official site: https://commoncrawl.org/
- The Pile paper (Gao et al., 2020): https://arxiv.org/abs/2101.00027
- RedPajama repository: https://github.com/togethercomputer/RedPajama-Data
- RefinedWeb paper (Penedo et al., 2023): https://arxiv.org/abs/2306.01116
- CulturaX paper (Nguyen et al., 2023): https://arxiv.org/abs/2309.09400
- Stanford Alpaca repository: https://github.com/tatsu-lab/stanford_alpaca
- UltraChat paper (Ding et al., 2023): https://arxiv.org/abs/2305.14233
- OpenAssistant dataset: https://huggingface.co/datasets/OpenAssistant/oasst1
- MMLU repository: https://github.com/hendrycks/test
- GSM8K paper (Cobbe et al., 2021): https://arxiv.org/abs/2110.14168
- HumanEval paper (Chen et al., 2021): https://arxiv.org/abs/2107.03374
- BBH paper (Suzgun et al., 2022): https://arxiv.org/abs/2210.09261
- HELM (Stanford CRFM): https://crfm.stanford.edu/helm/
- TruthfulQA repository: https://github.com/sylinrl/TruthfulQA
- LongBench paper (Bai et al., 2023): https://arxiv.org/abs/2308.14508
- C-Eval repository: https://github.com/SJTU-LIT/ceval
- datatrove repository: https://github.com/huggingface/datatrove