Theme
Evaluation & Benchmarks
Evaluation is using controlled methods to measure whether a language model "works" — it's the factual basis for all model decisions (whether to deploy, whether to switch the foundation, whether fine-tuning helped). Without evaluation, there's no engineering judgment — "feels better" isn't evidence, and "what's the leaderboard score" is only partial evidence.
This article first explains why LLM evaluation is hard, then unfolds three evaluation system categories, followed by capability-layered evaluation, contamination and leaderboard reading, and finally the offline/online evaluation loop. Hands-on methods are in Evaluation in Practice, and common benchmark scale and archive details at Datasets & Benchmarks.
1. Why LLM Evaluation Is So Hard
Compared to traditional machine learning (with fixed labels, accuracy calculable), LLM evaluation has four fundamental challenges:
- Generation tasks have no single answer: the same question can have multiple correct expressions; exact string matching will inevitably misjudge; but "loose matching" lets wrong answers slip through.
- Capabilities are highly subjective: whether a response is "helpful," "safe," or "has the right style" depends on human judgment, and standards differ across people.
- Capability boundaries are blurry: the model's capabilities change with scale, prompting method, and decoding parameters — what's evaluated may be only "performance under a certain configuration."
- Data contamination: evaluation questions may already appear in the pretraining corpus, inflating scores without the model knowing.
Three qualifiers for every evaluation result
Any benchmark score must answer three questions before comparison is valid: which model version, which prompt/sampling configuration, and has contamination been checked? For the same model, changing temperature from 0 to 0.8 can shift scores significantly; switching prompt wording can reverse leaderboard ranking. An isolated "XX score" means nothing.
2. Three Evaluation System Categories
LLM evaluation splits into three categories: intrinsic metrics, task benchmarks, and human & model judges. They answer different questions, cost different amounts, and mature teams use all three.
1. Intrinsic Metrics: How Well Does the Language Model Itself Fit
These only measure "how well the model fits the text distribution," without involving any downstream task. The core metric is perplexity:
text
Perplexity = exp( - (1/N) · Σ log P( w_t | w_<t ) )
= 2^cross-entropy
Intuition: the model's average "hesitation" per word when predicting an "unseen text."
Lower perplexity means the model is better at predicting text from that distribution.Perplexity's value and limits: it doesn't require labels, can rapidly compare "who fits the same distribution better," but it can't measure instruction following, safety, or factual correctness — a model with very low perplexity on historical corpus might completely ignore instructions. It's mainly used for pretraining stage monitoring (loss curves) and same-distribution corpus modeling comparison, detailed in Language Modeling.
2. Task Benchmarks: Giving the Model an "Exam"
Task benchmarks break capability into scorable question sets, using automatic metrics or controlled formats (multiple choice/limited answers) for evaluation. This is currently the most mainstream evaluation format:
| Benchmark | Year | What it tests | Format | Scale |
|---|---|---|---|---|
| MMLU | 2020 | Multi-domain knowledge across 57 disciplines (humanities/social sciences/sciences/medicine, etc.) | 4-option | ~16K questions |
| GSM8K | 2021 | Elementary school math word problems (multi-step reasoning) | Open-ended numeric | Test ~1.3K questions |
| MATH | 2021 | Competition-level math (AMC/AIME level) | Open-ended | 5K questions |
| HumanEval | 2021 | Code generation (function completion) | pass@k | 164 questions |
| MBPP | 2021 | Basic Python programming | pass@k | 974 questions |
| BBH | 2022 | 23 reasoning tasks from BIG-Bench where models don't beat humans | Multiple | 23 tasks |
| TruthfulQA | 2021 | Honesty: spotting common misconceptions and false beliefs | Q&A | 817 questions |
| ARC | 2018 | Science QA (elementary to middle school) | Choice/open | ~7.7K questions |
| HELM | 2022 | "Comprehensive health check" framework across multiple scenarios and metrics | Mixed | 42 scenarios × 7 metrics |
These benchmarks' automatic scoring mechanisms differ widely — before reading scores, understand the metrics:
| Metric | Meaning | Typical Use |
|---|---|---|
| Accuracy (multiple choice) | Option-level correctness | MMLU, ARC |
| Exact Match | Character-by-character match with reference answer | GSM8K numeric answers |
| F1 | Token overlap between generation and reference | Extractive QA |
| pass@k | Proportion of at least one correct answer among k generated candidates | HumanEval, MBPP |
| Robustness | Whether scores hold after question rewriting/perturbation | HELM's stress testing |
Taking HumanEval's pass@k as an example: code questions have no single answer; pass@1 is "one-shot generation passes tests"; pass@100 is "at least one of 100 generated candidates runs through" — the latter measures the model's "probability of generating correct code," closer to real coding-assist scenarios.
Detailed archive (scale, use, license) for these benchmarks is at Datasets & Benchmarks.
How to read benchmarks: a single benchmark only tests one side. GPT-4 era public data can serve as reference (e.g., MMLU ~86, GSM8K ~92, HumanEval ~67% pass@1, GPT-4o improved to ~90%, per official reports), but more important is understanding the four pitfalls of benchmark scores (see below).
Capability-layered profiles are more useful than single scores
More valuable than "total score" is a profile broken down by capability dimension: knowledge (MMLU), math (GSM8K/MATH), code (HumanEval/MBPP), reasoning (BBH), multilingual, long context (LongBench), tool use. A "high total score but weak code" model is unusable for a pure-code team.
3. Human and Model Judges: Measuring "Subjective Experience"
Task benchmarks test "is it right?" but product experience more often asks "is it good?" Three main types:
Human blind A/B: have humans rank (A/B) or score two responses. Gold standard, but expensive, slow, and annotator inter-rater agreement (Kappa) must be measured separately.
LLM-as-a-judge (model as judge): have a strong model (typically GPT-4 level) score or rank responses. The 2023 MT-Bench paper reported: GPT-4 judges align with human judgments at ~80% consistency, comparable to "human vs human" consistency — this is the factual basis for "model judges" scaling up.
Crowdsourced arena (Elo): like Chatbot Arena, users anonymously vote on two models' responses to the same question, ranked by Elo. Real users, real distribution, but unlabeled samples and uneven coverage.
Known biases of LLM-as-a-judge (must be guarded against when using):
| Bias | Manifestation | Mitigation |
|---|---|---|
| Position bias | Prefers responses listed first | Run twice with swapped order, average |
| Length bias | Longer responses score higher, regardless of quality | Explicitly constrain length, regularize |
| Self-preference bias | Judge prefers responses "from the same source" | Cross-validate with multiple judge models |
| Authority/format bias | Prefers well-formatted, confidently-worded responses | Score item by item by rubric rather than holistic score |
3. Capability-Layered Evaluation: Different Methods for Different Capabilities
Model capabilities aren't monolithic; evaluation must be designed by capability layer to pinpoint problems:
| Capability Layer | Typical Benchmark/Method | Key point |
|---|---|---|
| Language modeling foundation | Perplexity, loss curves | Use for pretraining stage monitoring |
| Knowledge | MMLU, ARC, subject-specific question banks | Watch for knowledge cutoff and contamination |
| Reasoning | GSM8K, MATH, BBH | Can pair with CoT prompting (see Prompting) |
| Code | HumanEval, MBPP | pass@k to assess generation diversity |
| Multilingual | MMLU multilingual version, translation benchmarks | Test Chinese/small-language separately |
| Long context | LongBench, needle-in-a-haystack | Link with Context & Long Context |
| Tool use / Agent | Tool call success rate, multi-step task completion | Chain evaluation |
| Honesty | TruthfulQA, hallucination detection | See Hallucination: Causes & Mitigation |
| Safety | Red team tests, jailbreak detection | See Safety & Risks |
The relationship between eval design and scaling laws
Capability-layering isn't arbitrary: a model's weaknesses usually follow scaling laws — some capabilities (like complex reasoning) emerge non-linearly as models grow, while others (like simple knowledge) just rise smoothly. Understanding the relationship between capability layers and scaling laws helps judge "whether this problem is solved by switching models or is a prompting/architecture-level issue."
4. Data Contamination and How to Read Leaderboards
1. Benchmark Contamination: The #1 Source of Inflated Scores
Benchmark contamination means evaluation questions appeared in the model's training corpus. LLMs crawl the entire web for pretraining, and many benchmarks (especially those released before 2021) have questions that circulated online — likely "memorized."
Empirical signals: models approach perfect scores on "seen" questions but drop sharply on "same-distribution but unseen variants"; or the model can verbatim reproduce questions from training corpus. GPT-4's technical report explicitly acknowledged "cannot fully rule out evaluation data contamination" and publicly shared some decontamination methods.
Leaderboards aren't facts — they're part of the test conditions
Seeing "XX model scores 95 on MMLU," first ask three questions: what prompting did it use? Was there a decontamination process? Did it publish variant-set scores simultaneously? The industry has widely seen "new-question live tests significantly below leaderboard numbers" — leaderboard scores are the ceiling, not the norm.
2. How to Read Leaderboards Correctly
- Compare under identical conditions: only comparable under the same eval framework and same sampling config; cross-framework score comparison is meaningless.
- Look at capability profiles, not total score: total scores mask weaknesses.
- Look at timestamps: a 2023 MMLU score can't be directly extrapolated to 2025 new-question live tests.
- Look for third-party replication: when official self-reporting vs community replication (like lm-eval-harness, OpenCompass results) disagree, trust the reproducible.
- Beware of "eval set overfitting": teams repeatedly tuning models on their own eval set effectively treat the eval set as training data, detaching scores from real capability — this is one of the classic ten pitfalls in Common Pitfalls & Anti-Patterns.
3. What a Credible Eval Report Should Disclose
A trustworthy eval report should disclose at least eight items — leaderboards missing any are suspect:
| Disclosure Item | Why It Matters |
|---|---|
| Model version and weight fingerprint | Version drift makes scores untraceable |
| Prompt and sampling config | Temperature, top-p, few-shot examples directly affect scores |
| Question set and splits | Which version, which subsets |
| Decontamination process | Whether n-gram overlap detection and manual spot-check were done |
| Eval framework version | Framework updates change scoring logic |
| Run count and variance | Sampling generation tasks need multiple averages |
| Compute resources | Affects resource-sensitive evals like long context |
| Reproduction entry | Config, code, results publicly reproducible? |
Treat this checklist as "the acceptance standard for evaluations" — it filters most marketing-style scores.
5. Offline and Online Evaluation: Evaluation Must Close the Loop
1. Self-Built Eval Sets: Born from Business
Public benchmarks test general capability, but what determines product success is the business scenario. Mature teams all build their own eval sets:
text
Five steps for self-built eval sets:
1. Sample real requests: randomly draw from online logs covering main scenarios
2. Define standards: first write "what counts as a good answer" (scoring rubric)
3. Build by category: split subsets by capability (knowledge/reasoning/format/safety)
4. Rotate regularly: prevent eval set overfitting (model "memorizing" the eval set)
5. Regression-run scores: run every model/prompt/param change through it firstSelf-built sets don't need to be large — dozens to hundreds of samples, the value lies in representativeness and consistent annotation standards. It's the bridge connecting "general benchmarks" with "business goals."
Evaluation isn't a "do once before launch" action; it's a closed loop running through the model lifecycle:
| Stage | Format | Purpose | Typical methods |
|---|---|---|---|
| Training | Offline automatic | Monitor loss, mid-run checkpoints | Perplexity, sample generation observation |
| Post-training | Offline benchmarks | Capability/safety regression | MMLU/GSM8K/red team |
| Pre-release | Offline human | Quality gate | Blind review, A/B, manual verification |
| Post-launch | Online metrics | Real-world effect | A/B testing, user feedback, review rate, satisfaction |
2. Eval Tools and Ecosystem
No need to build from scratch — open-source tools are mature: lm-eval-harness (EleutherAI) covers standard implementations of hundreds of benchmarks; OpenCompass (Shanghai AI Lab) supports multi-model Chinese/English comparison and leaderboard publishing; HELM (Stanford) provides multi-scenario multi-metric frameworks. Tool selection and hands-on are in Evaluation in Practice and Frameworks & Tool Selection.
Offline vs online, which weighs more? Offline is daily, online is the final verdict: offline eval (automatic benchmarks + offline human) covers controlled, cheap, fast — the daily workhorse; online eval (A/B, user behavior, manual annotation feedback) measures "whether users really think it's good in the real world" — the final arbiter. The gap between the two — "high offline scores, no online feeling" — is the most common eval failure mode, rooted in eval sets disconnected from the real distribution.
From evaluation to iteration loop
Best practice: build a golden set — dozens to hundreds of fixed samples covering key business scenarios; run it as regression before any fine-tuning/prompt change. Supplement new questions regularly outside the golden set to prevent overfitting. Full engineeringized approach (including LLM-as-a-judge implementation and bias handling) is in Evaluation in Practice.
Further Reading
- Hallucination: Causes & Mitigation — the connection between honesty evaluation and hallucination detection
- Safety & Risks — what evaluation tests for safety alignment
- Scaling Laws — the scaling patterns behind capability layering
- Evaluation in Practice — complete engineering from benchmarks to golden set
- Datasets & Benchmarks — archives of MMLU/GSM8K/HumanEval and more
- Common Pitfalls & Anti-Patterns — two major traps: "worshipping leaderboards" and "eval set overfitting"
References
- Hendrycks et al. Measuring Massive Multitask Language Understanding (MMLU, ICLR 2021) — the original MMLU paper
- Cobbe et al. Training Verifiers to Solve Math Word Problems (GSM8K, 2021) — the original GSM8K paper
- Chen et al. Evaluating Large Language Models Trained on Code (HumanEval, 2021) — HumanEval and Codex
- Zheng et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (NeurIPS 2023) — LLM-as-a-judge and Arena methodology
- Liang et al. Holistic Evaluation of Language Models (HELM, TMLR 2023) — multi-scenario multi-metric eval framework
- Suzgun et al. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them (BBH, 2022) — the original BBH paper
- OpenAI. GPT-4 Technical Report (2023) — official explanation of evaluation methods and contamination issues