Skip to content

Evaluation & Benchmarks

At a glance Evaluation is the only way to answer 'whether the model really works.' This article explains why LLM evaluation is hard, three categories of evaluation systems (intrinsic metrics/task benchmarks/human & model judges), capability-layered evaluation, benchmark contamination issues and how to read leaderboards correctly, and how offline and online evaluation form a closed loop.

Evaluation & Benchmarks ​

Evaluation is using controlled methods to measure whether a language model "works" — it's the factual basis for all model decisions (whether to deploy, whether to switch the foundation, whether fine-tuning helped). Without evaluation, there's no engineering judgment — "feels better" isn't evidence, and "what's the leaderboard score" is only partial evidence.

This article first explains why LLM evaluation is hard, then unfolds three evaluation system categories, followed by capability-layered evaluation, contamination and leaderboard reading, and finally the offline/online evaluation loop. Hands-on methods are in Evaluation in Practice, and common benchmark scale and archive details at Datasets & Benchmarks.

1. Why LLM Evaluation Is So Hard ​

Compared to traditional machine learning (with fixed labels, accuracy calculable), LLM evaluation has four fundamental challenges:

  1. Generation tasks have no single answer: the same question can have multiple correct expressions; exact string matching will inevitably misjudge; but "loose matching" lets wrong answers slip through.
  2. Capabilities are highly subjective: whether a response is "helpful," "safe," or "has the right style" depends on human judgment, and standards differ across people.
  3. Capability boundaries are blurry: the model's capabilities change with scale, prompting method, and decoding parameters — what's evaluated may be only "performance under a certain configuration."
  4. Data contamination: evaluation questions may already appear in the pretraining corpus, inflating scores without the model knowing.

Three qualifiers for every evaluation result

Any benchmark score must answer three questions before comparison is valid: which model version, which prompt/sampling configuration, and has contamination been checked? For the same model, changing temperature from 0 to 0.8 can shift scores significantly; switching prompt wording can reverse leaderboard ranking. An isolated "XX score" means nothing.

2. Three Evaluation System Categories ​

LLM evaluation splits into three categories: intrinsic metrics, task benchmarks, and human & model judges. They answer different questions, cost different amounts, and mature teams use all three.

1. Intrinsic Metrics: How Well Does the Language Model Itself Fit ​

These only measure "how well the model fits the text distribution," without involving any downstream task. The core metric is perplexity:

text
Perplexity = exp( - (1/N) · Σ log P( w_t | w_<t ) )
             = 2^cross-entropy

Intuition: the model's average "hesitation" per word when predicting an "unseen text."
Lower perplexity means the model is better at predicting text from that distribution.

Perplexity's value and limits: it doesn't require labels, can rapidly compare "who fits the same distribution better," but it can't measure instruction following, safety, or factual correctness — a model with very low perplexity on historical corpus might completely ignore instructions. It's mainly used for pretraining stage monitoring (loss curves) and same-distribution corpus modeling comparison, detailed in Language Modeling.

2. Task Benchmarks: Giving the Model an "Exam" ​

Task benchmarks break capability into scorable question sets, using automatic metrics or controlled formats (multiple choice/limited answers) for evaluation. This is currently the most mainstream evaluation format:

BenchmarkYearWhat it testsFormatScale
MMLU2020Multi-domain knowledge across 57 disciplines (humanities/social sciences/sciences/medicine, etc.)4-option~16K questions
GSM8K2021Elementary school math word problems (multi-step reasoning)Open-ended numericTest ~1.3K questions
MATH2021Competition-level math (AMC/AIME level)Open-ended5K questions
HumanEval2021Code generation (function completion)pass@k164 questions
MBPP2021Basic Python programmingpass@k974 questions
BBH202223 reasoning tasks from BIG-Bench where models don't beat humansMultiple23 tasks
TruthfulQA2021Honesty: spotting common misconceptions and false beliefsQ&A817 questions
ARC2018Science QA (elementary to middle school)Choice/open~7.7K questions
HELM2022"Comprehensive health check" framework across multiple scenarios and metricsMixed42 scenarios × 7 metrics

These benchmarks' automatic scoring mechanisms differ widely — before reading scores, understand the metrics:

MetricMeaningTypical Use
Accuracy (multiple choice)Option-level correctnessMMLU, ARC
Exact MatchCharacter-by-character match with reference answerGSM8K numeric answers
F1Token overlap between generation and referenceExtractive QA
pass@kProportion of at least one correct answer among k generated candidatesHumanEval, MBPP
RobustnessWhether scores hold after question rewriting/perturbationHELM's stress testing

Taking HumanEval's pass@k as an example: code questions have no single answer; pass@1 is "one-shot generation passes tests"; pass@100 is "at least one of 100 generated candidates runs through" — the latter measures the model's "probability of generating correct code," closer to real coding-assist scenarios.

Detailed archive (scale, use, license) for these benchmarks is at Datasets & Benchmarks.

How to read benchmarks: a single benchmark only tests one side. GPT-4 era public data can serve as reference (e.g., MMLU ~86, GSM8K ~92, HumanEval ~67% pass@1, GPT-4o improved to ~90%, per official reports), but more important is understanding the four pitfalls of benchmark scores (see below).

Capability-layered profiles are more useful than single scores

More valuable than "total score" is a profile broken down by capability dimension: knowledge (MMLU), math (GSM8K/MATH), code (HumanEval/MBPP), reasoning (BBH), multilingual, long context (LongBench), tool use. A "high total score but weak code" model is unusable for a pure-code team.

3. Human and Model Judges: Measuring "Subjective Experience" ​

Task benchmarks test "is it right?" but product experience more often asks "is it good?" Three main types:

Human blind A/B: have humans rank (A/B) or score two responses. Gold standard, but expensive, slow, and annotator inter-rater agreement (Kappa) must be measured separately.

LLM-as-a-judge (model as judge): have a strong model (typically GPT-4 level) score or rank responses. The 2023 MT-Bench paper reported: GPT-4 judges align with human judgments at ~80% consistency, comparable to "human vs human" consistency — this is the factual basis for "model judges" scaling up.

Crowdsourced arena (Elo): like Chatbot Arena, users anonymously vote on two models' responses to the same question, ranked by Elo. Real users, real distribution, but unlabeled samples and uneven coverage.

Known biases of LLM-as-a-judge (must be guarded against when using):

BiasManifestationMitigation
Position biasPrefers responses listed firstRun twice with swapped order, average
Length biasLonger responses score higher, regardless of qualityExplicitly constrain length, regularize
Self-preference biasJudge prefers responses "from the same source"Cross-validate with multiple judge models
Authority/format biasPrefers well-formatted, confidently-worded responsesScore item by item by rubric rather than holistic score

3. Capability-Layered Evaluation: Different Methods for Different Capabilities ​

Model capabilities aren't monolithic; evaluation must be designed by capability layer to pinpoint problems:

Capability LayerTypical Benchmark/MethodKey point
Language modeling foundationPerplexity, loss curvesUse for pretraining stage monitoring
KnowledgeMMLU, ARC, subject-specific question banksWatch for knowledge cutoff and contamination
ReasoningGSM8K, MATH, BBHCan pair with CoT prompting (see Prompting)
CodeHumanEval, MBPPpass@k to assess generation diversity
MultilingualMMLU multilingual version, translation benchmarksTest Chinese/small-language separately
Long contextLongBench, needle-in-a-haystackLink with Context & Long Context
Tool use / AgentTool call success rate, multi-step task completionChain evaluation
HonestyTruthfulQA, hallucination detectionSee Hallucination: Causes & Mitigation
SafetyRed team tests, jailbreak detectionSee Safety & Risks

The relationship between eval design and scaling laws

Capability-layering isn't arbitrary: a model's weaknesses usually follow scaling laws — some capabilities (like complex reasoning) emerge non-linearly as models grow, while others (like simple knowledge) just rise smoothly. Understanding the relationship between capability layers and scaling laws helps judge "whether this problem is solved by switching models or is a prompting/architecture-level issue."

4. Data Contamination and How to Read Leaderboards ​

1. Benchmark Contamination: The #1 Source of Inflated Scores ​

Benchmark contamination means evaluation questions appeared in the model's training corpus. LLMs crawl the entire web for pretraining, and many benchmarks (especially those released before 2021) have questions that circulated online — likely "memorized."

Empirical signals: models approach perfect scores on "seen" questions but drop sharply on "same-distribution but unseen variants"; or the model can verbatim reproduce questions from training corpus. GPT-4's technical report explicitly acknowledged "cannot fully rule out evaluation data contamination" and publicly shared some decontamination methods.

Leaderboards aren't facts — they're part of the test conditions

Seeing "XX model scores 95 on MMLU," first ask three questions: what prompting did it use? Was there a decontamination process? Did it publish variant-set scores simultaneously? The industry has widely seen "new-question live tests significantly below leaderboard numbers" — leaderboard scores are the ceiling, not the norm.

2. How to Read Leaderboards Correctly ​

  • Compare under identical conditions: only comparable under the same eval framework and same sampling config; cross-framework score comparison is meaningless.
  • Look at capability profiles, not total score: total scores mask weaknesses.
  • Look at timestamps: a 2023 MMLU score can't be directly extrapolated to 2025 new-question live tests.
  • Look for third-party replication: when official self-reporting vs community replication (like lm-eval-harness, OpenCompass results) disagree, trust the reproducible.
  • Beware of "eval set overfitting": teams repeatedly tuning models on their own eval set effectively treat the eval set as training data, detaching scores from real capability — this is one of the classic ten pitfalls in Common Pitfalls & Anti-Patterns.

3. What a Credible Eval Report Should Disclose ​

A trustworthy eval report should disclose at least eight items — leaderboards missing any are suspect:

Disclosure ItemWhy It Matters
Model version and weight fingerprintVersion drift makes scores untraceable
Prompt and sampling configTemperature, top-p, few-shot examples directly affect scores
Question set and splitsWhich version, which subsets
Decontamination processWhether n-gram overlap detection and manual spot-check were done
Eval framework versionFramework updates change scoring logic
Run count and varianceSampling generation tasks need multiple averages
Compute resourcesAffects resource-sensitive evals like long context
Reproduction entryConfig, code, results publicly reproducible?

Treat this checklist as "the acceptance standard for evaluations" — it filters most marketing-style scores.

5. Offline and Online Evaluation: Evaluation Must Close the Loop ​

1. Self-Built Eval Sets: Born from Business ​

Public benchmarks test general capability, but what determines product success is the business scenario. Mature teams all build their own eval sets:

text
Five steps for self-built eval sets:
1. Sample real requests: randomly draw from online logs covering main scenarios
2. Define standards: first write "what counts as a good answer" (scoring rubric)
3. Build by category: split subsets by capability (knowledge/reasoning/format/safety)
4. Rotate regularly: prevent eval set overfitting (model "memorizing" the eval set)
5. Regression-run scores: run every model/prompt/param change through it first

Self-built sets don't need to be large — dozens to hundreds of samples, the value lies in representativeness and consistent annotation standards. It's the bridge connecting "general benchmarks" with "business goals."

Evaluation isn't a "do once before launch" action; it's a closed loop running through the model lifecycle:

StageFormatPurposeTypical methods
TrainingOffline automaticMonitor loss, mid-run checkpointsPerplexity, sample generation observation
Post-trainingOffline benchmarksCapability/safety regressionMMLU/GSM8K/red team
Pre-releaseOffline humanQuality gateBlind review, A/B, manual verification
Post-launchOnline metricsReal-world effectA/B testing, user feedback, review rate, satisfaction

2. Eval Tools and Ecosystem ​

No need to build from scratch — open-source tools are mature: lm-eval-harness (EleutherAI) covers standard implementations of hundreds of benchmarks; OpenCompass (Shanghai AI Lab) supports multi-model Chinese/English comparison and leaderboard publishing; HELM (Stanford) provides multi-scenario multi-metric frameworks. Tool selection and hands-on are in Evaluation in Practice and Frameworks & Tool Selection.

Offline vs online, which weighs more? Offline is daily, online is the final verdict: offline eval (automatic benchmarks + offline human) covers controlled, cheap, fast — the daily workhorse; online eval (A/B, user behavior, manual annotation feedback) measures "whether users really think it's good in the real world" — the final arbiter. The gap between the two — "high offline scores, no online feeling" — is the most common eval failure mode, rooted in eval sets disconnected from the real distribution.

From evaluation to iteration loop

Best practice: build a golden set — dozens to hundreds of fixed samples covering key business scenarios; run it as regression before any fine-tuning/prompt change. Supplement new questions regularly outside the golden set to prevent overfitting. Full engineeringized approach (including LLM-as-a-judge implementation and bias handling) is in Evaluation in Practice.

Further Reading ​

References ​