Skip to content

Evaluation in Practice

At a glance Build a layered evaluation system combining 'public benchmarks + custom golden sets + LLM-as-a-judge + regression testing + online evaluation': with batch evaluation code, judge implementation and bias control, and cost management — turning 'is the model good?' into quantifiable, regressible engineering metrics.

Evaluation in Practice ​

Evaluation is the measure and standard of LLM engineering: without it, all prompt tuning, fine-tuning, and RAG optimization are just feelings; with it, you gain the engineering capability of "improvements are attributable, regressions are detectable, and launches are confident."

Why is LLM evaluation hard? Because outputs are open-ended text rather than fixed labels, and dimensions like factuality, safety, and style cannot be captured by a single score. This article provides a layered practical system: public benchmarks answer "what is the model's absolute level," custom sets answer "how does it perform in your scenario," judges and humans answer "what is the quality and safety," and regression testing answers "whether changes broke old capabilities." Theoretical foundations are at Evaluation and Benchmarks.

I. Four-Layer Evaluation Architecture ​

LayerQuestion AnsweredMethodCostFrequency
① Intrinsic metricsHow is the model's "fluency / certainty"?perplexity, lossLowDuring training, in real time
② Task benchmarksHow are general capabilities (knowledge / math / code)?MMLU / GSM8K / HumanEval, etc.MediumPer model version
③ Custom set + judge/humanHow does the model perform on business tasks?golden set + scoringMedium–HighPer change
④ Online evaluationHow do real users feel?A/B tests, feedback, telemetryHighContinuous

Most "evaluation systems that are ready for launch" refer to ②③④. This article focuses on these three layers.

Evaluation Metrics: Fix the "What Counts as a Score" First ​

Before building evaluation, fix "how we count" — different metrics yield different conclusions:

MetricDefinitionSuitable For
Accuracy (exact match)Proportion of outputs matching the ground truth character-by-characterTasks with ground truth (classification / extraction)
Partial / substring matchKeywords or substrings of the ground truth appear in outputSummarization, open-ended Q&A
pass@kProportion of at least one pass within k generationsCode generation
Human score (1–5)Annotators score by rubricQuality, style, safety
Judge scoreLLM scores by rubricSame as above but automated
Win rate (pairwise)Proportion of wins in two-model comparisonRanking, selection

The discipline of fixing metrics: once a metric is chosen, it must be used consistently across version comparisons. Switching from accuracy to judge scores is like changing rulers — the numbers before and after are not comparable. This is the prerequisite for regression testing to work.

II. Benchmark Sets + Custom Sets: Walking on Two Legs ​

1. Public Benchmarks: Pick Meaningful Subsets ​

Public benchmarks (MMLU, GSM8K, HumanEval, BBH, HELM, etc., full archive at Datasets and Benchmarks Archive) are suitable for horizontal comparison and regression, but running the full sets is expensive. Engineering advice:

BenchmarkMeasuresCommon Practice
MMLUMulti-domain knowledgeSample subsets by discipline (e.g., 20 questions from each of 10 disciplines) for quick regression
GSM8KMath reasoningSample 100–200 questions (answers can be auto-compared, suitable for CI)
HumanEvalCode generationFull 164 questions is fine, pass@1 auto-scoring
Custom golden setYour businessSee below
bash
# Run a model with lm-evaluation-harness (ready to use; note: HF access required)
pip install lm_eval

lm_eval --model hf \
  --model_args pretrained=Qwen/Qwen2.5-7B-Instruct \
  --tasks mmlu,gsm8k \
  --batch_size 4 \
  --output_path results/ \
  --log_samples

Benchmark scores ≠ your scores

Benchmark scores are heavily influenced by implementation details (number of few-shot examples, prompt templates, decoding parameters, tokenizer versions). The same model can differ by several points across different harnesses. Fix your evaluation configuration and note it — otherwise comparisons like "we scored 0.73, they scored 0.75" are meaningless. More pitfalls of leaderboard obsession are at Common Pitfalls and Anti-Patterns.

2. Custom Sets: Golden Sets Are the Core of the Evaluation System ​

A golden set (gold-standard set) is a fixed dataset collected from real business scenarios with ground-truth answers / evaluation criteria. Three construction disciplines:

DisciplineDescription
Real sourcesCollect from online logs, customer service records, and real queries — don't make them up
Authoritative answersGround-truth answers annotated by domain experts, or "correct/incorrect" criteria
Append, never modifyAdd new cases to the golden set; never delete or modify old cases (otherwise regression testing becomes distorted)
json
// Example golden set entry (task: customer service intent classification)
[
  {
    "id": "G-0001",
    "input": "Your app keeps crashing, I'm so frustrated!",
    "expected": {"intent": "complaint", "label": "complaint"},
    "note": "Boundary: strong tone but still a complaint, not a refund inquiry"
  },
  {
    "id": "G-0002",
    "input": "Can I still use it if my membership card has expired?",
    "expected": {"intent": "inquiry", "label": "inquiry"}
  }
]

III. LLM-as-a-Judge: Implementation and Bias ​

1. When to Use Judges vs. Humans ​

ScenarioPrimary Choice
Ground truth available (classification, extraction, code, math)Rule-based / auto-scoring, no judge needed
Open-ended quality (summarization quality, answer relevance, politeness)LLM-as-a-judge
High-risk content (safety, factual disputes, legal/medical)Human annotation (judges can assist screening)

2. Judge Implementation ​

Three judge paradigms (in order of reliability): direct scoring → rubric scoring → pairwise comparison. Pairwise comparison is usually the most stable.

python
# Judge scoring example: rubric mode (much more reliable than "give 1–10")
JUDGE_PROMPT = """You are an evaluator. Score the answer based on the following criteria (1–5):
- 5: Fully addresses the question, accurate information, clear structure
- 3: Partially relevant, with minor redundancy or bias
- 1: Irrelevant or clearly wrong
Output only a number.

[Question] {question}
[Reference Answer] {reference}
[Answer to Evaluate] {candidate}
Score:"""

def judge_score(question, reference, candidate):
    prompt = JUDGE_PROMPT.format(question=question, reference=reference, candidate=candidate)
    out = call_judge_model(prompt)   # recommend using a stronger model for the judge
    return int(out.strip()[:1])
python
# Pairwise comparison: let the judge choose between A and B, count win rate (most robust ranking signal)
def judge_compare(question, answer_a, answer_b):
    prompt = f"""Compare the two answers below for usefulness, accuracy, and completeness. Which is better?
[Question] {question}
[Answer A] {answer_a}
[Answer B] {answer_b}
Output only "A" or "B":"""
    return call_judge_model(prompt).strip()

3. Judge Bias and Mitigation ​

BiasManifestationMitigation
Position biasAnswers listed first are always preferredSwap A/B order and score both ways; flag disagreements for human review
Self-preferenceJudge prefers answers "like itself" (same family models)Use models from different families as judges; or manually review sensitive samples
Length biasLonger answers score higherAdd "redundancy deduction" to rubric; keep comparison lengths similar
Over-generous scoringJudge tends to score too highUse pairwise mode + regular manual calibration sampling

Where judges cannot replace humans

Factual errors are the easiest for judges to miss — they may share the same training data and hallucination patterns as the model being evaluated. Factual evaluation either uses golden sets with answers or requires manual spot-checking of samples the judge rated "excellent." Mechanisms and evaluation of hallucination are at Hallucination: Causes and Mitigation.

IV. Offline Batch Evaluation: A Reusable Pipeline ​

Turn evaluation into scripted batch processing, not "manual one-by-one testing":

python
import json, random
from concurrent.futures import ThreadPoolExecutor

def evaluate(model_fn, golden_set, judge_fn=None, n=50):
    """Sample from golden set to run the model, supporting rule-based or judge scoring"""
    random.seed(42)
    sample = random.sample(golden_set, min(n, len(golden_set)))
    results = []
    for case in sample:
        out = model_fn(case["input"])
        verdict = rule_check(out, case["expected"])          # Rule-based check
        # Or: verdict = judge_fn(case["input"], case.get("reference"), out)  # Judge scoring
        results.append({**case, "output": out, "verdict": verdict})
    return results

def summarize(results):
    """Output summary: accuracy + failed samples (for error analysis)"""
    ok = [r for r in results if r["verdict"] is True]
    errors = [r for r in results if r["verdict"] is not True]
    print(f"Pass rate: {len(ok)/len(results):.1%}  ({len(ok)}/{len(results)})")
    return errors

# Usage: run the same function on every change, compare two summarize outputs
r1 = evaluate(my_model_v1, golden_set)
r2 = evaluate(my_model_v2, golden_set)

Three engineering disciplines:

  1. Fix random seeds and model call configurations (results are reproducible at temperature=0).
  2. Output failed samples separately: half the value of evaluation lies in "error analysis" — look at where the model fails and whether the failure mode is patterned.
  3. Persist results: record each evaluation's time, version, and scores in a table/database to form trend curves.

Error Analysis: Turning Scores into Actions ​

Half the value of evaluation lies in "error analysis." After getting failure samples, go through the checklist systematically instead of "glancing and moving on":

ObservationLikely ConclusionNext Action
Errors concentrate on a specific input typeSystematic issue with prompts/data for that typeWrite targeted test cases and fixes for that class
Almost all are format errorsInsufficient output constraintsApply JSON mode / constrained decoding
Almost all are factual errorsInsufficient knowledge or hallucinationRetrieval-augmented generation (RAG), add fact checks
Randomly distributed, no patternProbabilistic noise or sample size too smallIncrease sampling rounds, enlarge eval set
Strongly correlated with model versionBehavior drift in a specific versionLocate the change, rollback or compensate

Error analysis produces a "known issues list" (issue type → impact scope → remediation direction), which serves both as R&D backlog input and as pre-launch "known limitations" disclosure.

V. Regression Testing and Golden Sets ​

Regression testing = using a fixed golden set + fixed judge/rules to auto-run every change, preventing "fixing A breaks B." This is a key step in LLM application engineering:

yaml
# Pseudocode: regression gate in CI
# 1. Code/prompt/model weights change → trigger regression
# 2. Run three categories of tests:
#    - Unit tests: structured output is valid (JSON schema validation)
#    - Golden regression: business accuracy no lower than baseline (e.g., 95% baseline)
#    - Safety regression: failure rate on dangerous/violating inputs meets threshold
# 3. All pass → allow merge / launch

Three layers of regression testing

Unit tests (format / field validity, lowest cost) → golden regression (business accuracy, medium cost) → safety/fact spot-checks (high risk, high cost). Ensure layer one is all green first, then gradually add layers two and three. Once regression gates are established, you can "keep changing prompts" without fear of breaking things.

Multi-Model Horizontal Comparison: The Standard Action for Selection ​

"The right model to pick" is the most commonly asked evaluation question. The standard approach is to turn selection into an evaluation experiment:

Candidate model list (including base, quantized, distilled versions)
    │
    ▼
One golden set + one set of prompts + identical decoding parameters
    │
    ▼
Run metrics for each model + costs (token price × avg output length)
    │
    ▼
Decide on three dimensions: quality / latency / cost

Key discipline: prompts and decoding parameters must be identical during comparison, otherwise you're testing prompts rather than models. Cost and latency dimensions are covered in the accounting methods at Deployment and Serving.

Evaluate "quantized versions" in selection experiments, not just full precision

Production often uses INT4 quantized models. If selection only tests fp16 but production uses AWQ, your evaluation results will mismatch production behavior. Selection experiments should test the exact tier you plan to launch (quantization, context length, batch parameters all aligned).

Evaluation Reports: Numbers Must Be Retellable ​

An evaluation's output should be a retellable report, not a screen of logs:

FieldContent
VersionModel / prompt / data version (git-commit level)
ConfigEval set, judge model, decoding parameters
ResultsScores per metric + sample size
TrendComparison with previous run / baseline
Error analysisFailure type distribution and representative cases
ConclusionWhether to switch, next steps

Reports should be saved as markdown or entered into tables, so any team member can retell "why this version was chosen" — this is the ultimate test of whether an evaluation system is trusted.

VI. Online Evaluation: The Final Arbiter Is the User ​

Offline evaluation cannot cover the diversity of real inputs. Four essentials for online evaluation:

MethodPracticeMetric
A/B testingNew version to 5–10% of traffic only, compare with oldTask completion rate, click-through, conversion
Manual spot-checksSample 1%–5% of conversations for annotator scoringSatisfaction, error rate
User feedbackThumbs up/down buttons, "report problem" entryNegative feedback rate
TelemetryResponse latency, failure rate, retry rateStability

Ethics and compliance of online evaluation

Online evaluation involving real user data (especially when generated content is shown to users) must comply with data regulations and product ethics. High-risk output scenarios (legal, medical, financial) must have human review fallback; see Safety and Risks for discussion.

VII. Evaluation Cost Management ​

Evaluation is also "compute overhead" and needs budget management:

MethodDescription
Layered samplingWhen the golden set is large, only run 30–50 random samples each time; save the full set for pre-release
Cascading strategyRule-based check first (free), fall to judge (costs money) if it fails, fall to human if judge fails again
Reuse resultsCache input-output-score; don't re-call for the same input
Use cheaper judgesUse mini models for routine tasks, strong models for high-risk sample review
Evaluation frequency tieringUnit tests every run, golden daily, full benchmarks weekly

The positive cycle between evaluation and cost

Evaluation seems like "extra overhead," but it prevents far more expensive mistakes: one undiscovered regression incident's customer service / PR cost far exceeds the compute cost of running a few thousand evals. Treat the evaluation budget as insurance, not a cost center.

Further Reading ​

References ​