Theme
Evaluation in Practice
In one sentence: Evaluation is the process of turning "model performance" into "actionable business evidence" — a good evaluation doesn't just answer "what score the model got," but also "is this score credible, where did it go wrong, and is it worth deploying."
A single accuracy on an offline test set is just the starting point. This article walks through metric selection, all the way to online evaluation, covering evaluation design, statistical rigor, and failure analysis. Theoretical definitions of evaluation metrics are in Deep Learning Evaluation and Experiments, and experiment discipline is in DL Design Principles.
1. Metric Selection: Ask the Business First
Metrics must be determined by business costs, not by defaulting to accuracy. Three typical scenarios:
| Scenario | Common metrics | Why |
|---|---|---|
| Classification (imbalanced classes) | precision / recall / F1 / PR-AUC | Accuracy is drowned by the majority class; cost of false negatives vs false positives is typically asymmetric |
| Classification (multi-class) | macro-F1, per-class report, confusion matrix | See each class individually, not just an average |
| Regression | MAE / RMSE / Huber loss | RMSE is sensitive to large errors (is that what your business wants?); MAE is robust |
| Ranking / recommendation | NDCG@k, HitRate@k, MAP | Business cares about "does the top-k hit?" (see Deep Learning Recommender Systems) |
| Generative | BLEU/ROUGE + human evaluation | Automated metrics correlate poorly with human judgment; use a dual-track approach |
| Binary classification + adjustable threshold | ROC-AUC, PR-AUC | Threshold selection is a business decision (sensitivity vs specificity), not a model decision |
An unintuitive pitfall
Don't use accuracy as the primary metric when classes are imbalanced. When 99% are negative, predicting all-negative gives 99% accuracy but zero business value. Medical screening, fraud detection, and search ranking are almost always these scenarios — at that point, precision/recall curves and business cost functions are the meaningful benchmarks.
2. Confidence Intervals and Multiple Runs
Deep learning has randomness (initialization, data order, augmentation sampling), and a single run's score is a random variable, not a conclusion.
2.1 Standard Approach: Multiple Runs
Fix a hyperparameter configuration, run N times (N ≥ 5) with different random seeds, and report mean ± standard deviation:
python
import numpy as np
from scipy import stats
def run_experiment(seed):
set_seed(seed)
model = train() # Full training
return evaluate(model) # Return test set metric
accs = [run_experiment(seed) for seed in range(5)]
mean, std = np.mean(accs), np.std(accs, ddof=1)
# 95% confidence interval (t-distribution, more accurate for small samples)
ci = stats.t.interval(0.95, df=len(accs) - 1, loc=mean, scale=std / np.sqrt(len(accs)))
print(f"acc = {mean:.4f} ± {std:.4f} (95% CI: {ci[0]:.4f} ~ {ci[1]:.4f})")2.2 When the Error Is Smaller Than the Difference, It Doesn't Matter
Model A has 90.0% accuracy and Model B has 90.1% — if their respective standard deviations are ±0.3%, that 0.1-point difference is within noise, and you can't claim B is better. The correct approach:
- Paired testing: Compare the two models in pairs using the same test samples and the same seeds (e.g., McNemar's test for classification error counts).
- Report error bars: Always attach confidence intervals to any key comparison, letting readers judge whether the difference is outside noise.
3. Evaluation Set Construction Discipline: Preventing Leakage
Evaluation results come in two flavors: credible and self-deceptive. Leakage is the most common source of self-deception (mechanics elaborated in Common Pitfalls and Anti-Patterns):
| Leakage Type | Example | Consequence |
|---|---|---|
| Data leakage | Normalization mean/std computed using full data (including test set) | Inflated test metrics |
| Temporal leakage | Using future news to predict yesterday's prices | Production collapse |
| Identity leakage | Different photos from the same user appear in both train/test | Unreliable metrics |
| Augmentation leakage | Augmentation random seeds shared between train and test | Irreproducible results |
| Tuning leakage | Repeatedly using the test set to pick models | Test set becomes validation set |
Construction discipline:
- Split order: The first thing you do with data is split into train / val / test (and fix it). After that, only touch train and val.
- Time-series data: Split by time, never shuffle; consider "lagged validation" (use a historical window to predict a future window, see Data and Data Engineering).
- Group homogeneous samples: All samples from the same entity (user/patient/session) must go into the same partition to prevent identity leakage.
- Touch the test set only once: Hyperparameter tuning, model selection, and early stopping all happen on the validation set. The test set is used only once at the final checkpoint.
4. Failure Analysis and Error Set Review
Metrics answer "what score," and failure analysis answers "why points were deducted." Steps:
python
# Collect metadata for all misclassified samples (prediction, ground truth, confidence, source)
errors = []
model.eval()
with torch.no_grad():
for x, y, meta in test_loader: # Assume dataset returns (x, y, meta)
logits = model(x)
probs = torch.softmax(logits, dim=-1)
pred = probs.argmax(-1)
for i in range(len(y)):
if pred[i] != y[i]:
errors.append({
"true": y[i].item(), "pred": pred[i].item(),
"conf": probs[i].max().item(), "meta": meta[i],
})
# Analyze patterns:
# 1. Group by class: which classes confuse each other? (confusion matrix)
# 2. Group by confidence: high-confidence errors = model is "confidently wrong," the most dangerous
# 3. Group by source: do errors concentrate in a particular data batch/time period/group?Standard actions for error set review:
- Plot the confusion matrix to see systematic confusion patterns (e.g., "4 vs 9," "cat vs dog").
- Bucket error rates: Slice by confidence, by duration, by demographic — find the model's blind spots.
- Manually read 50 error samples: Label them as "reasonably wrong" (humans would also struggle) vs "unreasonably wrong" (model/data problem).
- Errors → Actions: Fix data (labeling errors), add data (blind spot classes), modify the model (mistaking background for landmarks, see Interpretability and Fairness).
5. A/B Testing and Online Evaluation
There's often a gap between offline metrics and online performance (distribution shift, latency, interaction effects). The gold standard before deployment is A/B testing:
| Stage | Question | Method |
|---|---|---|
| Offline | Are model candidates good? | Test set + confidence intervals |
| Shadow mode | How does the new model perform online? | Run the new model alongside production, log only, no decisions |
| Small-traffic A/B | Do business metrics improve? | Route 5%–10%, run for 1–2 weeks |
| Full rollout | Long-term effect | Monitor drift and have a rollback switch |
A/B testing key points:
- Predefine success metrics and minimum effect (e.g., conversion rate up 0.5% with a confidence interval excluding 0). Avoid "storytelling with data."
- Watch out for Simpson's paradox: Overall no improvement but every subgroup improves (or vice versa) — always analyze by strata (new vs returning users, region, etc.).
- Pay attention to the long tail: A model with good average metrics but degraded p99 shouldn't ship.
- Have a rollback switch: If the new model breaks something, be able to revert with one click.
Continuous online evaluation involves monitoring, drift detection, and model update cadence — that falls under MLOps and Model Deployment.
6. Human Evaluation for Generative Models
Automated metrics for generative models (text, image, audio) like BLEU, FID, and Inception Score correlate poorly with human perception. You must supplement with structured human evaluation (basics in Generative Models and Diffusion Models and Generative AI):
- Dimensional scoring: Don't rate "good or bad"; score by dimensions (relevance, fluency, factual accuracy, harmfulness, each on a 1–5 scale). Define scoring criteria, use multiple raters, and report inter-rater agreement (e.g., Cohen's kappa).
- Paired preference: Ask raters to choose one (Model A vs Model B output). This is more stable than absolute scores and most closely resembles A/B testing.
- Error taxonomy: Categorize generation errors (hallucination, repetition, instruction deviation), track proportions — this is more actionable than a single score.
- Scale and cost: Human evaluation is expensive and slow; use reasonable sampling (~200–500 per config) plus LLM-based pre-screening (see Large Language Models (LLM)).
7. Case Study: Complete Image Classification Evaluation Pipeline
Bringing it all together into runnable code:
python
import numpy as np
import torch, torch.nn as nn
from torch.utils.data import DataLoader
from sklearn.metrics import (confusion_matrix, classification_report,
precision_recall_fscore_support)
import matplotlib.pyplot as plt
def full_evaluation(model, loader, device):
"""Returns (probs, preds, labels) and prints a full evaluation report"""
model.eval()
probs_list, preds_list, labels_list = [], [], []
with torch.no_grad():
for x, y in loader:
logits = model(x.to(device))
probs = torch.softmax(logits, dim=-1)
probs_list.append(probs.cpu().numpy())
preds_list.append(probs.argmax(-1).cpu().numpy())
labels_list.append(y.numpy())
probs = np.concatenate(probs_list)
preds = np.concatenate(preds_list)
labels = np.concatenate(labels_list)
return probs, preds, labels
def evaluate_with_report(model, test_loader, classes, device, n_runs=5):
"""Multiple runs + confidence intervals + per-class metrics + confusion matrix + error analysis"""
all_run_accs = []
for seed in range(n_runs):
torch.manual_seed(seed)
# Note: same seed only re-runs randomness during evaluation; training randomness is handled in the experiment script
probs, preds, labels = full_evaluation(model, test_loader, device)
all_run_accs.append((preds == labels).mean())
if seed == 0:
print("=== per-class report ===")
print(classification_report(labels, preds, target_names=classes, digits=4))
cm = confusion_matrix(labels, preds)
plt.imshow(cm, cmap="Blues")
plt.colorbar()
plt.title("Confusion matrix")
plt.savefig("confusion_matrix.png")
mean_acc = np.mean(all_run_accs)
std_acc = np.std(all_run_accs, ddof=1)
print(f"\naccuracy over {n_runs} runs: {mean_acc:.4f} ± {std_acc:.4f}")
# Error analysis: high-confidence misclassifications ("confidently wrong")
wrong = preds != labels
conf_wrong = probs[wrong].max(axis=-1)
top_idx = np.argsort(conf_wrong)[-10:] # Top 10 highest-confidence errors
print("top confident mistakes (index, true, pred, conf):")
for i in top_idx:
print(f" {i}: true={labels[i]}, pred={preds[i]}, conf={conf_wrong[i]:.3f}")
return mean_acc, std_acc, cmKey usage notes:
- Use the same
test_loaderdefinition for evaluation and training (same transforms, same data). Otherwise, you're not evaluating the model that was trained. - Average across multiple runs with mean ± std to make error bars like "±0.3%" visible.
- Per-class report directly exposes weaknesses in imbalanced classes.
- High-confidence error list is the seed for error analysis — manually inspecting these 10 samples often reveals data issues or model blind spots.
The artifacts from this pipeline (scores, confidence intervals, confusion matrix, error samples) form the evidence package for whether a model is "ready for deployment," and they're also the material for the "Results" section of Portfolio Projects.
Further Reading
- Deep Learning Evaluation and Experiments — theoretical definitions and statistical foundations for metrics
- Common Pitfalls and Anti-Patterns — complete symptom list for evaluation leakage
- Building a Deep Learning Project from Scratch — evaluation as a project artifact
- MLOps and Model Deployment — online evaluation and monitoring
- Interpretability and Fairness — downstream of error analysis
- Datasets and Tools Reference — data sources for evaluation set construction
References
- Dietterich. Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms (Neural Computation 1998) — Paired tests and classifier comparison
- Kohavi et al. Online Controlled Experiments at Large Scale (KDD 2013) — A/B testing engineering practice
- scikit-learn. Model evaluation: quantifying the quality of predictions — Authoritative reference for metrics API