Theme
Evaluation in Practice
Evaluation is the measure and standard of LLM engineering: without it, all prompt tuning, fine-tuning, and RAG optimization are just feelings; with it, you gain the engineering capability of "improvements are attributable, regressions are detectable, and launches are confident."
Why is LLM evaluation hard? Because outputs are open-ended text rather than fixed labels, and dimensions like factuality, safety, and style cannot be captured by a single score. This article provides a layered practical system: public benchmarks answer "what is the model's absolute level," custom sets answer "how does it perform in your scenario," judges and humans answer "what is the quality and safety," and regression testing answers "whether changes broke old capabilities." Theoretical foundations are at Evaluation and Benchmarks.
I. Four-Layer Evaluation Architecture
| Layer | Question Answered | Method | Cost | Frequency |
|---|---|---|---|---|
| ① Intrinsic metrics | How is the model's "fluency / certainty"? | perplexity, loss | Low | During training, in real time |
| ② Task benchmarks | How are general capabilities (knowledge / math / code)? | MMLU / GSM8K / HumanEval, etc. | Medium | Per model version |
| ③ Custom set + judge/human | How does the model perform on business tasks? | golden set + scoring | Medium–High | Per change |
| ④ Online evaluation | How do real users feel? | A/B tests, feedback, telemetry | High | Continuous |
Most "evaluation systems that are ready for launch" refer to ②③④. This article focuses on these three layers.
Evaluation Metrics: Fix the "What Counts as a Score" First
Before building evaluation, fix "how we count" — different metrics yield different conclusions:
| Metric | Definition | Suitable For |
|---|---|---|
| Accuracy (exact match) | Proportion of outputs matching the ground truth character-by-character | Tasks with ground truth (classification / extraction) |
| Partial / substring match | Keywords or substrings of the ground truth appear in output | Summarization, open-ended Q&A |
| pass@k | Proportion of at least one pass within k generations | Code generation |
| Human score (1–5) | Annotators score by rubric | Quality, style, safety |
| Judge score | LLM scores by rubric | Same as above but automated |
| Win rate (pairwise) | Proportion of wins in two-model comparison | Ranking, selection |
The discipline of fixing metrics: once a metric is chosen, it must be used consistently across version comparisons. Switching from accuracy to judge scores is like changing rulers — the numbers before and after are not comparable. This is the prerequisite for regression testing to work.
II. Benchmark Sets + Custom Sets: Walking on Two Legs
1. Public Benchmarks: Pick Meaningful Subsets
Public benchmarks (MMLU, GSM8K, HumanEval, BBH, HELM, etc., full archive at Datasets and Benchmarks Archive) are suitable for horizontal comparison and regression, but running the full sets is expensive. Engineering advice:
| Benchmark | Measures | Common Practice |
|---|---|---|
| MMLU | Multi-domain knowledge | Sample subsets by discipline (e.g., 20 questions from each of 10 disciplines) for quick regression |
| GSM8K | Math reasoning | Sample 100–200 questions (answers can be auto-compared, suitable for CI) |
| HumanEval | Code generation | Full 164 questions is fine, pass@1 auto-scoring |
| Custom golden set | Your business | See below |
bash
# Run a model with lm-evaluation-harness (ready to use; note: HF access required)
pip install lm_eval
lm_eval --model hf \
--model_args pretrained=Qwen/Qwen2.5-7B-Instruct \
--tasks mmlu,gsm8k \
--batch_size 4 \
--output_path results/ \
--log_samplesBenchmark scores ≠ your scores
Benchmark scores are heavily influenced by implementation details (number of few-shot examples, prompt templates, decoding parameters, tokenizer versions). The same model can differ by several points across different harnesses. Fix your evaluation configuration and note it — otherwise comparisons like "we scored 0.73, they scored 0.75" are meaningless. More pitfalls of leaderboard obsession are at Common Pitfalls and Anti-Patterns.
2. Custom Sets: Golden Sets Are the Core of the Evaluation System
A golden set (gold-standard set) is a fixed dataset collected from real business scenarios with ground-truth answers / evaluation criteria. Three construction disciplines:
| Discipline | Description |
|---|---|
| Real sources | Collect from online logs, customer service records, and real queries — don't make them up |
| Authoritative answers | Ground-truth answers annotated by domain experts, or "correct/incorrect" criteria |
| Append, never modify | Add new cases to the golden set; never delete or modify old cases (otherwise regression testing becomes distorted) |
json
// Example golden set entry (task: customer service intent classification)
[
{
"id": "G-0001",
"input": "Your app keeps crashing, I'm so frustrated!",
"expected": {"intent": "complaint", "label": "complaint"},
"note": "Boundary: strong tone but still a complaint, not a refund inquiry"
},
{
"id": "G-0002",
"input": "Can I still use it if my membership card has expired?",
"expected": {"intent": "inquiry", "label": "inquiry"}
}
]III. LLM-as-a-Judge: Implementation and Bias
1. When to Use Judges vs. Humans
| Scenario | Primary Choice |
|---|---|
| Ground truth available (classification, extraction, code, math) | Rule-based / auto-scoring, no judge needed |
| Open-ended quality (summarization quality, answer relevance, politeness) | LLM-as-a-judge |
| High-risk content (safety, factual disputes, legal/medical) | Human annotation (judges can assist screening) |
2. Judge Implementation
Three judge paradigms (in order of reliability): direct scoring → rubric scoring → pairwise comparison. Pairwise comparison is usually the most stable.
python
# Judge scoring example: rubric mode (much more reliable than "give 1–10")
JUDGE_PROMPT = """You are an evaluator. Score the answer based on the following criteria (1–5):
- 5: Fully addresses the question, accurate information, clear structure
- 3: Partially relevant, with minor redundancy or bias
- 1: Irrelevant or clearly wrong
Output only a number.
[Question] {question}
[Reference Answer] {reference}
[Answer to Evaluate] {candidate}
Score:"""
def judge_score(question, reference, candidate):
prompt = JUDGE_PROMPT.format(question=question, reference=reference, candidate=candidate)
out = call_judge_model(prompt) # recommend using a stronger model for the judge
return int(out.strip()[:1])python
# Pairwise comparison: let the judge choose between A and B, count win rate (most robust ranking signal)
def judge_compare(question, answer_a, answer_b):
prompt = f"""Compare the two answers below for usefulness, accuracy, and completeness. Which is better?
[Question] {question}
[Answer A] {answer_a}
[Answer B] {answer_b}
Output only "A" or "B":"""
return call_judge_model(prompt).strip()3. Judge Bias and Mitigation
| Bias | Manifestation | Mitigation |
|---|---|---|
| Position bias | Answers listed first are always preferred | Swap A/B order and score both ways; flag disagreements for human review |
| Self-preference | Judge prefers answers "like itself" (same family models) | Use models from different families as judges; or manually review sensitive samples |
| Length bias | Longer answers score higher | Add "redundancy deduction" to rubric; keep comparison lengths similar |
| Over-generous scoring | Judge tends to score too high | Use pairwise mode + regular manual calibration sampling |
Where judges cannot replace humans
Factual errors are the easiest for judges to miss — they may share the same training data and hallucination patterns as the model being evaluated. Factual evaluation either uses golden sets with answers or requires manual spot-checking of samples the judge rated "excellent." Mechanisms and evaluation of hallucination are at Hallucination: Causes and Mitigation.
IV. Offline Batch Evaluation: A Reusable Pipeline
Turn evaluation into scripted batch processing, not "manual one-by-one testing":
python
import json, random
from concurrent.futures import ThreadPoolExecutor
def evaluate(model_fn, golden_set, judge_fn=None, n=50):
"""Sample from golden set to run the model, supporting rule-based or judge scoring"""
random.seed(42)
sample = random.sample(golden_set, min(n, len(golden_set)))
results = []
for case in sample:
out = model_fn(case["input"])
verdict = rule_check(out, case["expected"]) # Rule-based check
# Or: verdict = judge_fn(case["input"], case.get("reference"), out) # Judge scoring
results.append({**case, "output": out, "verdict": verdict})
return results
def summarize(results):
"""Output summary: accuracy + failed samples (for error analysis)"""
ok = [r for r in results if r["verdict"] is True]
errors = [r for r in results if r["verdict"] is not True]
print(f"Pass rate: {len(ok)/len(results):.1%} ({len(ok)}/{len(results)})")
return errors
# Usage: run the same function on every change, compare two summarize outputs
r1 = evaluate(my_model_v1, golden_set)
r2 = evaluate(my_model_v2, golden_set)Three engineering disciplines:
- Fix random seeds and model call configurations (results are reproducible at temperature=0).
- Output failed samples separately: half the value of evaluation lies in "error analysis" — look at where the model fails and whether the failure mode is patterned.
- Persist results: record each evaluation's time, version, and scores in a table/database to form trend curves.
Error Analysis: Turning Scores into Actions
Half the value of evaluation lies in "error analysis." After getting failure samples, go through the checklist systematically instead of "glancing and moving on":
| Observation | Likely Conclusion | Next Action |
|---|---|---|
| Errors concentrate on a specific input type | Systematic issue with prompts/data for that type | Write targeted test cases and fixes for that class |
| Almost all are format errors | Insufficient output constraints | Apply JSON mode / constrained decoding |
| Almost all are factual errors | Insufficient knowledge or hallucination | Retrieval-augmented generation (RAG), add fact checks |
| Randomly distributed, no pattern | Probabilistic noise or sample size too small | Increase sampling rounds, enlarge eval set |
| Strongly correlated with model version | Behavior drift in a specific version | Locate the change, rollback or compensate |
Error analysis produces a "known issues list" (issue type → impact scope → remediation direction), which serves both as R&D backlog input and as pre-launch "known limitations" disclosure.
V. Regression Testing and Golden Sets
Regression testing = using a fixed golden set + fixed judge/rules to auto-run every change, preventing "fixing A breaks B." This is a key step in LLM application engineering:
yaml
# Pseudocode: regression gate in CI
# 1. Code/prompt/model weights change → trigger regression
# 2. Run three categories of tests:
# - Unit tests: structured output is valid (JSON schema validation)
# - Golden regression: business accuracy no lower than baseline (e.g., 95% baseline)
# - Safety regression: failure rate on dangerous/violating inputs meets threshold
# 3. All pass → allow merge / launchThree layers of regression testing
Unit tests (format / field validity, lowest cost) → golden regression (business accuracy, medium cost) → safety/fact spot-checks (high risk, high cost). Ensure layer one is all green first, then gradually add layers two and three. Once regression gates are established, you can "keep changing prompts" without fear of breaking things.
Multi-Model Horizontal Comparison: The Standard Action for Selection
"The right model to pick" is the most commonly asked evaluation question. The standard approach is to turn selection into an evaluation experiment:
Candidate model list (including base, quantized, distilled versions)
│
▼
One golden set + one set of prompts + identical decoding parameters
│
▼
Run metrics for each model + costs (token price × avg output length)
│
▼
Decide on three dimensions: quality / latency / costKey discipline: prompts and decoding parameters must be identical during comparison, otherwise you're testing prompts rather than models. Cost and latency dimensions are covered in the accounting methods at Deployment and Serving.
Evaluate "quantized versions" in selection experiments, not just full precision
Production often uses INT4 quantized models. If selection only tests fp16 but production uses AWQ, your evaluation results will mismatch production behavior. Selection experiments should test the exact tier you plan to launch (quantization, context length, batch parameters all aligned).
Evaluation Reports: Numbers Must Be Retellable
An evaluation's output should be a retellable report, not a screen of logs:
| Field | Content |
|---|---|
| Version | Model / prompt / data version (git-commit level) |
| Config | Eval set, judge model, decoding parameters |
| Results | Scores per metric + sample size |
| Trend | Comparison with previous run / baseline |
| Error analysis | Failure type distribution and representative cases |
| Conclusion | Whether to switch, next steps |
Reports should be saved as markdown or entered into tables, so any team member can retell "why this version was chosen" — this is the ultimate test of whether an evaluation system is trusted.
VI. Online Evaluation: The Final Arbiter Is the User
Offline evaluation cannot cover the diversity of real inputs. Four essentials for online evaluation:
| Method | Practice | Metric |
|---|---|---|
| A/B testing | New version to 5–10% of traffic only, compare with old | Task completion rate, click-through, conversion |
| Manual spot-checks | Sample 1%–5% of conversations for annotator scoring | Satisfaction, error rate |
| User feedback | Thumbs up/down buttons, "report problem" entry | Negative feedback rate |
| Telemetry | Response latency, failure rate, retry rate | Stability |
Ethics and compliance of online evaluation
Online evaluation involving real user data (especially when generated content is shown to users) must comply with data regulations and product ethics. High-risk output scenarios (legal, medical, financial) must have human review fallback; see Safety and Risks for discussion.
VII. Evaluation Cost Management
Evaluation is also "compute overhead" and needs budget management:
| Method | Description |
|---|---|
| Layered sampling | When the golden set is large, only run 30–50 random samples each time; save the full set for pre-release |
| Cascading strategy | Rule-based check first (free), fall to judge (costs money) if it fails, fall to human if judge fails again |
| Reuse results | Cache input-output-score; don't re-call for the same input |
| Use cheaper judges | Use mini models for routine tasks, strong models for high-risk sample review |
| Evaluation frequency tiering | Unit tests every run, golden daily, full benchmarks weekly |
The positive cycle between evaluation and cost
Evaluation seems like "extra overhead," but it prevents far more expensive mistakes: one undiscovered regression incident's customer service / PR cost far exceeds the compute cost of running a few thousand evals. Treat the evaluation budget as insurance, not a cost center.
Further Reading
- Evaluation and Benchmarks — Theoretical treatment of three evaluation types, benchmark contamination, and how to read leaderboards
- Hallucination: Causes and Mitigation — The factual dimension most easily missed by judge evaluation
- Fine-Tuning in Practice: Full LoRA Workflow — Concrete before/after comparison cases for fine-tuning
- Prompting in Practice — Baseline → hypothesis → regression methodology for prompt optimization
- Datasets and Benchmarks Archive — Full profiles of benchmarks like MMLU/GSM8K/HumanEval
- Common Pitfalls and Anti-Patterns — Leaderboard obsession, evaluation overfitting, data contamination, and other evaluation pitfalls
- Glossary — Quick reference for terms like golden set, judge, pass@k
References
- Measuring Massive Multitask Language Understanding (MMLU, arXiv:2009.03300) — Original MMLU paper
- Training Verifiers to Solve Math Word Problems (GSM8K, arXiv:2110.14168) — Original GSM8K paper
- Evaluating Large Language Models Trained on Code (HumanEval, arXiv:2107.03374) — Original HumanEval paper
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv:2306.05685) — Systematic study of LLM-as-a-judge bias
- lm-evaluation-harness (GitHub) — EleutherAI open-source evaluation framework
- OpenCompass (GitHub) — Shanghai AI Lab open-source evaluation platform
- OpenAI Evals (GitHub) — OpenAI open-source evaluation framework