Skip to content

Datasets and Benchmarks Reference

At a glance A comprehensive reference of data across the entire LLM pipeline: pretraining corpora, post-training instruction data, and mainstream evaluation benchmarks — including scale, purpose, how to access, and licensing considerations, from Common Crawl to MT-Bench and C-Eval.

This page contains time-sensitive content, current as of 2025-08; job descriptions, rankings, product features, and other information may have changed. Please verify with original sources before citing.

Datasets and Benchmarks Reference ​

This page is a one-stop reference for data and evaluation benchmarks in the LLM world: pretraining corpora (building the model), post-training instruction data (teaching the model), and evaluation benchmarks (measuring the model). Each entry includes scale, construction method, use case, a one-line description, and licensing notes. The scale and leaderboard numbers for data and benchmarks are continuously updated — this page's data is current as of August 2025; always refer to official sources for the latest figures.

Three paths through this page

Working on pretraining / continued pretraining → see "I. Pretraining Corpora"; Working on fine-tuning / alignment → see "II. Post-training and Instruction Data"; Working on evaluation / model selection → see "III. Evaluation Benchmarks" and "IV. Chinese and Multilingual Data." Data sourcing and cleaning details are covered in Pretraining: Data and Objectives.

I. Pretraining Corpora ​

Pretraining corpora are "the model's textbooks": scale and quality directly determine the model's capability ceiling. Internet crawl data (Common Crawl) is the raw material for most corpora and must be cleaned, deduplicated, filtered, and blended before it's usable.

CorpusScaleConstruction methodOne-line descriptionLicensing notes
Common Crawl~3–5 billion webpages per month, PB-scaleFull-internet crawling on a monthly basis since 2008A freely available web crawl corpus — the "raw ore" for all large-scale pretraining corporaData is publicly downloadable, but page content is copyrighted by the originating websites; respect robots.txt and copyright
The Pile~825 GiB across 22 subsetsHand-curated mix of books, papers, code, web pages, etc. by EleutherAIThe first "carefully blended" open pretraining corpus, used by models like GPT-NeoXLicenses vary by subset; includes copyrighted text (e.g. books), so commercial use requires per-subset evaluation
RedPajama-Data-1T~1.2 trillion tokensTOGETHER's recreation of the Llama corpus blendAn open reproduction of Llama's 7-way data blend (67% web / 4.5% code / 4.5% books, etc.)Permissive and commercially usable; both source code and data are open
RefinedWebPublic version: ~600B tokensCommon Crawl cleaned by TII (the Falcon team)A pure-web corpus emphasizing "high-quality cleaning + rigorous deduplication," used by the Falcon seriesBased on Common Crawl; follow the same copyright requirements as the source
CulturaX~6.3 trillion tokens, 167 languagesCleaned and deduplicated merge of mC4 and OSCARA multilingual corpus covering 167 languages with strong coverage of low-resource non-English languagesCopyright status for multilingual data is complex; evaluate target-language corpora before commercial use

1. From Raw Crawl to Training Corpus: the Cleaning Pipeline ​

Raw Common Crawl is extremely noisy (ads, garbled text, duplicates, low-quality pages) and cannot be used for training directly. The standard pipeline has four steps:

  1. Collection and parsing: Download WARC snapshots, extract body text (strip HTML tags, navigation, ads), and determine language per URL.
  2. Quality filtering: Remove low-quality pages using rule-based criteria (length, punctuation density, symbol ratio) and heuristic scores (language model perplexity, classifiers). Also filter for toxicity and privacy.
  3. Deduplication: MinHash approximate deduplication + exact deduplication to remove duplicated content across pages (repeated data wastes compute and causes the model to "memorize test answers").
  4. Blending and sampling: Mix corpora according to multilingual/code/math/book ratios, then up-sample or down-sample by perplexity.

Why everyone starts from Common Crawl

Common Crawl is free, massive, and continuously updated — it's the "raw material" of data. But the real difference between corpora lies in the cleaning pipeline: deduplication algorithms, quality scorers, and language blends all vary. "Data is the model" mainly refers to these processing details. New corpora (like CulturaX, RefinedWeb) are mostly "better cleaning" rather than "newer sources."

II. Post-training and Instruction Data ​

Post-training data is orders of magnitude smaller than pretraining corpora (millions vs. trillions of tokens), but quality and diversity determine how "obedient" the model becomes. Note: a large amount of instruction data is generated by closed commercial models, raising licensing and copyright concerns (see the warning box at the end).

DatasetScaleSourceOne-line descriptionLicensing notes
Alpaca~52K instruction–response pairsStanford, generated using text-davinci-003The classic open SFT dataset using self-instruct to expand instructionsGenerated by an OpenAI model; officially restricted to academic research use
ShareGPT~90K conversation turnsCommunity-shared ChatGPT conversation logsReal user conversations that served as the training source for models like VicunaGenerated by ChatGPT and shared by users — license is ambiguous, use with caution commercially
UltraChat~1.5M conversation turnsTsinghua University, generated via ChatGPT across 30 topicsLarge-scale multi-turn conversations covering instructions, queries, and chit-chatGenerated by a commercial model; primarily for research use
OpenAssistant (oasst1)~16K conversation treesLAION, crowdsourced human-writtenHigh-quality purely human-written multi-turn conversations — no "AI-generated" controversyOpen source (Apache 2.0), commercially usable
Dolly~15K entriesDatabricks employees, hand-writtenHuman-written instruction–response pairs, small but clean — great for testingOpen source and commercially usable

1. Three Ways to Produce Instruction Data ​

MethodExamplesProsCons
Human-writtenoasst1, DollyHigh quality, clean copyright, commercially usableHigh cost, limited scale
Model-generatedAlpaca, UltraChat, ShareGPTCheap, can scale up massivelyCopyright controversies; student models inherit the teacher model's flaws
Human-in-the-loopMost commercial SFT dataA balance between quality and scaleComplex workflow, still not cheap

2. Scale Reference Values ​

StageTypical scaleNotes
Pretraining~20–30B tokens for a 1B model; ~2–4T tokens for a 7B modelChinchilla recommends "approximately 1:20 ratio for parameters to data tokens"; frontier models generally use 10T+ tokens
SFT1K–100K high-quality instructions1,000 high-quality examples can already make a noticeable difference (LIMA experiment); quality and diversity are the keys
Preference alignment (RLHF/DPO)Tens of thousands to hundreds of thousands of preference pairsMust cover typical failure modes for helpfulness, honesty, and safety

Quality over quantity is the iron rule

For post-training data, "1,000 high-quality examples" often delivers more value than "100,000 low-quality filler examples." The scale references above are starting points only; your final call should be based on what works on your actual task. See Fine-tuning and Alignment for more.

Compliance red lines for AI-generated data

Alpaca, ShareGPT, UltraChat, and others are generated by OpenAI/Anthropic commercial models — OpenAI's terms of service explicitly prohibit "using outputs to train competing models," and the copyright status of AI-generated content is legally contested. Use only for research and internal experiments; for commercial products, prioritize human-written or licensed data. See Fine-tuning: SFT and PEFT for discussion on data quality and quantity.

III. Evaluation Benchmarks ​

Evaluation benchmarks are "model exam papers," classified by the capabilities they assess. Keep three things in mind: leaderboard scores are only comparable when using "the same prompts and implementation"; older benchmarks like MMLU may suffer from data contamination; and gaming a single benchmark ≠ real capability (see Evaluation and Benchmarks and Evaluation in Practice).

1. General Knowledge and Language Understanding ​

BenchmarkScaleOne-line description
GLUE9 tasks, ~250K samplesA 2018 text-understanding benchmark, the de facto standard of the BERT era, now largely saturated
SuperGLUE8 more difficult reasoning tasksGLUE's successor, released in 2019, also now largely "beaten" by LLMs
MMLU57 disciplines, ~16K multiple-choice questionsThe most classic general-knowledge benchmark, long the #1 leaderboard metric
HELM42 scenarios × 7 metricsStanford's multi-metric evaluation framework emphasizing "transparency" and fair comparisons

2. Math and Reasoning ​

BenchmarkScaleOne-line description
GSM8K~8K grade-school math word problemsElementary/middle school arithmetic, testing multi-step reasoning; used with CoT
MATH~12.5K competition-level problemsHigh school to competition-level math across 5 difficulty levels, with strong discrimination
BBH23 hard tasks from BIG-BenchTests logic, deduction, multi-step reasoning, and other "smart" tasks — small models generally can't solve them
ARC~7.8K science questionsAI2's science QA (ARC-Challenge is harder), testing common sense and scientific knowledge

3. Code ​

BenchmarkScaleOne-line description
HumanEval164 Python function completion tasksThe de facto benchmark for code generation, evaluated with pass@k
MBPP974 beginner-level Python tasksA more basic programming benchmark, often reported alongside HumanEval

4. Long Context and Conversations ​

BenchmarkScaleOne-line description
LongBench21 tasks, ~4,750 questionsBilingual (Chinese/English) long-context evaluation with average input lengths of 10K+ tokens, measuring "how long can the model actually read"
MT-Bench80 multi-turn open-ended questionsA conversational quality benchmark scored by LLM-as-a-judge (GPT-4); an important reference for Chatbot Arena
TruthfulQA817 questionsSpecifically designed to test "hallucination" and factual accuracy: whether the model produces plausible-sounding but incorrect answers

5. Instruction Following and More ​

BenchmarkScaleOne-line description
IFEval~500 instructionsSpecifically tests "whether the model actually follows instructions" (format/constraint adherence), with good discrimination
BIG-Bench204 tasksThe largest heterogeneous benchmark; in practice, the LLM community mostly uses the harder subset, BBH
SQuAD~100K questionsA 2016 extractive reading comprehension benchmark, a classic task of the BERT era

6. Four Common Pitfalls of Evaluation Benchmarks ​

  1. Data contamination: Benchmarks like MMLU have been publicly available for years, so questions may have leaked into training data, inflating scores. Defenses include deduplication, monitoring, and building private test sets.
  2. Saturation and score gaming: GLUE/SuperGLUE are fully saturated; MMLU is approaching saturation — when discrimination drops, the #1 and #10 ranked models may be in the same tier.
  3. Inconsistent evaluation methodology: Whether CoT prompting is used, how many few-shot examples are provided, and what temperature is set all significantly affect scores. Cross-model comparisons must use the same methodology.
  4. Leaderboards ≠ real business needs: Benchmarks measure "exam-taking ability." Real-world performance depends on your own task distribution — always build your own evaluation set.

How to choose benchmarks

  • General capability comparison: the trio of MMLU + GSM8K + HumanEval (low cost, reproducible);
  • Math and reasoning: GSM8K + MATH + BBH;
  • Code: HumanEval + MBPP;
  • Long context: LongBench (+ build your own "needle in a haystack" task);
  • Conversation and "real-world feel": MT-Bench / Chatbot Arena;
  • Chinese: see the next section on C-Eval / CMMLU. For building evaluations in practice, see Evaluation in Practice.

IV. Chinese and Multilingual Data ​

Chinese evaluation and data is a must-have for Chinese developers, so we list them in a dedicated table:

DatasetScaleUse caseOne-line description
C-Eval~14K questionsChinese general knowledge52 disciplines, covering middle school to graduate level — the Chinese MMLU
CMMLU~11.5K questionsChinese general knowledge67 disciplines including humanities and social science sub-categories; complements C-Eval
CLUE9 task typesChinese language understandingThe Chinese GLUE, the Chinese standard of the BERT era
SuperCLUEMulti-task suiteChinese conversational abilityA Chinese conversational benchmark with basic/advanced/expert difficulty tiers
BELLE~1M Chinese instructionsChinese SFTBaidu's open-source self-instruct Chinese instruction data
COIG~190K Chinese instructionsChinese SFTMulti-source Chinese instruction data released by BAAI (Beijing Academy of AI)

Multilingual best practices

For Chinese LLM evaluation, run at least C-Eval + CMMLU (knowledge) + Chinese code/conversation subsets, and check the model's tokenizer efficiency on Chinese — some Western-prioritized models consume significantly more tokens on Chinese input (see Tokenization and Vocabulary).

V. Licensing and Usage Considerations ​

Copyright and AI-generated data: compliance red lines

Pretraining corpora: Common Crawl and its derivatives containa large amount of copyrighted material, and several countries have seen lawsuits over training data. Commercial products must evaluate data sources and regional laws carefully. AI-generated instruction data: Alpaca, ShareGPT, UltraChat, and others are generated by OpenAI/Anthropic commercial models — OpenAI's terms prohibit "training competing models with outputs," and the copyright of generated content is legally contested. Use only for research and internal experiments. Evaluation benchmarks: Most benchmarks carry research licenses (academic use only); check each dataset's LICENSE before large-scale commercial evaluation.

A four-question compliance checklist

  1. Is the data source explicitly licensed for your intended use? 2. Does the service terms of any model used to generate the data prohibit using it for training? 3. What are the legal risks in your target jurisdiction (e.g., the EU AI Act, national copyright rulings)? 4. Has your commercial product completed a data-lineage audit? If all four are satisfied, you're clear to launch.

VI. Data Processing Toolchain ​

ToolUse caseLink
Hugging Face Datasets libraryStandard interface for dataset loading, streaming, splitting, and exporthttps://huggingface.co/docs/datasets
datatroveLarge-scale data cleaning pipeline (the tool that cleaned FineWeb) — deduplication, filtering, and scoring end-to-endhttps://github.com/huggingface/datatrove
lm-evaluation-harnessStandard framework for running evaluation benchmarks like MMLU, GSM8K, and HumanEvalhttps://github.com/EleutherAI/lm-evaluation-harness
OpenCompassChinese and multimodal evaluation suite with leaderboard serviceshttps://github.com/open-compass/opencompass

VII. Five Steps to Build a Data Pipeline from Scratch ​

A minimal viable workflow for readers who want to use their own data for pretraining or fine-tuning:

  1. Define goals and budget: Start with the model scale and training token budget (see the "scale reference values" above), then work backward to figure out how much raw data you need — cleaning and deduplication typically discard 50%–90% of raw data.
  2. Choose data sources: For pretraining, use open corpora (RedPajama/RefinedWeb) or a custom crawler; for instruction data, prioritize human-written or human-in-the-loop data, avoiding unauthorized AI-generated data (see "V. Licensing").
  3. Clean and deduplicate: Run quality filtering, MinHash deduplication, and toxicity filtering with datatrove; split multilingual corpora by language.
  4. Blend and sample: Decide on multilingual/code/math/book ratios based on your task objectives; up-sample "useful" corpora using perplexity.
  5. Validate quality: Start with a small-scale trial run using 100M tokens, observe loss curves and downstream benchmark (C-Eval/MMLU) changes across data-cleaning steps — every step of the data pipeline should be validated by experiments, not intuition.

Further Reading ​

References ​