Skip to content

Evaluation in Practice

Quick overview Accuracy isn't the whole story. This article covers evaluation design in depth: choosing metrics based on business context, confidence intervals and multiple runs, A/B testing and online evaluation, failure analysis and error set review, evaluation set construction with anti-leakage discipline, human evaluation for generative models, and runnable code for a complete image classification evaluation pipeline.

Evaluation in Practice ​

In one sentence: Evaluation is the process of turning "model performance" into "actionable business evidence" — a good evaluation doesn't just answer "what score the model got," but also "is this score credible, where did it go wrong, and is it worth deploying."

A single accuracy on an offline test set is just the starting point. This article walks through metric selection, all the way to online evaluation, covering evaluation design, statistical rigor, and failure analysis. Theoretical definitions of evaluation metrics are in Deep Learning Evaluation and Experiments, and experiment discipline is in DL Design Principles.

1. Metric Selection: Ask the Business First ​

Metrics must be determined by business costs, not by defaulting to accuracy. Three typical scenarios:

ScenarioCommon metricsWhy
Classification (imbalanced classes)precision / recall / F1 / PR-AUCAccuracy is drowned by the majority class; cost of false negatives vs false positives is typically asymmetric
Classification (multi-class)macro-F1, per-class report, confusion matrixSee each class individually, not just an average
RegressionMAE / RMSE / Huber lossRMSE is sensitive to large errors (is that what your business wants?); MAE is robust
Ranking / recommendationNDCG@k, HitRate@k, MAPBusiness cares about "does the top-k hit?" (see Deep Learning Recommender Systems)
GenerativeBLEU/ROUGE + human evaluationAutomated metrics correlate poorly with human judgment; use a dual-track approach
Binary classification + adjustable thresholdROC-AUC, PR-AUCThreshold selection is a business decision (sensitivity vs specificity), not a model decision

An unintuitive pitfall

Don't use accuracy as the primary metric when classes are imbalanced. When 99% are negative, predicting all-negative gives 99% accuracy but zero business value. Medical screening, fraud detection, and search ranking are almost always these scenarios — at that point, precision/recall curves and business cost functions are the meaningful benchmarks.

2. Confidence Intervals and Multiple Runs ​

Deep learning has randomness (initialization, data order, augmentation sampling), and a single run's score is a random variable, not a conclusion.

2.1 Standard Approach: Multiple Runs ​

Fix a hyperparameter configuration, run N times (N ≥ 5) with different random seeds, and report mean ± standard deviation:

python
import numpy as np
from scipy import stats

def run_experiment(seed):
    set_seed(seed)
    model = train()          # Full training
    return evaluate(model)   # Return test set metric

accs = [run_experiment(seed) for seed in range(5)]
mean, std = np.mean(accs), np.std(accs, ddof=1)

# 95% confidence interval (t-distribution, more accurate for small samples)
ci = stats.t.interval(0.95, df=len(accs) - 1, loc=mean, scale=std / np.sqrt(len(accs)))
print(f"acc = {mean:.4f} ± {std:.4f}  (95% CI: {ci[0]:.4f} ~ {ci[1]:.4f})")

2.2 When the Error Is Smaller Than the Difference, It Doesn't Matter ​

Model A has 90.0% accuracy and Model B has 90.1% — if their respective standard deviations are ±0.3%, that 0.1-point difference is within noise, and you can't claim B is better. The correct approach:

  • Paired testing: Compare the two models in pairs using the same test samples and the same seeds (e.g., McNemar's test for classification error counts).
  • Report error bars: Always attach confidence intervals to any key comparison, letting readers judge whether the difference is outside noise.

3. Evaluation Set Construction Discipline: Preventing Leakage ​

Evaluation results come in two flavors: credible and self-deceptive. Leakage is the most common source of self-deception (mechanics elaborated in Common Pitfalls and Anti-Patterns):

Leakage TypeExampleConsequence
Data leakageNormalization mean/std computed using full data (including test set)Inflated test metrics
Temporal leakageUsing future news to predict yesterday's pricesProduction collapse
Identity leakageDifferent photos from the same user appear in both train/testUnreliable metrics
Augmentation leakageAugmentation random seeds shared between train and testIrreproducible results
Tuning leakageRepeatedly using the test set to pick modelsTest set becomes validation set

Construction discipline:

  1. Split order: The first thing you do with data is split into train / val / test (and fix it). After that, only touch train and val.
  2. Time-series data: Split by time, never shuffle; consider "lagged validation" (use a historical window to predict a future window, see Data and Data Engineering).
  3. Group homogeneous samples: All samples from the same entity (user/patient/session) must go into the same partition to prevent identity leakage.
  4. Touch the test set only once: Hyperparameter tuning, model selection, and early stopping all happen on the validation set. The test set is used only once at the final checkpoint.

4. Failure Analysis and Error Set Review ​

Metrics answer "what score," and failure analysis answers "why points were deducted." Steps:

python
# Collect metadata for all misclassified samples (prediction, ground truth, confidence, source)
errors = []
model.eval()
with torch.no_grad():
    for x, y, meta in test_loader:        # Assume dataset returns (x, y, meta)
        logits = model(x)
        probs = torch.softmax(logits, dim=-1)
        pred = probs.argmax(-1)
        for i in range(len(y)):
            if pred[i] != y[i]:
                errors.append({
                    "true": y[i].item(), "pred": pred[i].item(),
                    "conf": probs[i].max().item(), "meta": meta[i],
                })

# Analyze patterns:
# 1. Group by class: which classes confuse each other? (confusion matrix)
# 2. Group by confidence: high-confidence errors = model is "confidently wrong," the most dangerous
# 3. Group by source: do errors concentrate in a particular data batch/time period/group?

Standard actions for error set review:

  • Plot the confusion matrix to see systematic confusion patterns (e.g., "4 vs 9," "cat vs dog").
  • Bucket error rates: Slice by confidence, by duration, by demographic — find the model's blind spots.
  • Manually read 50 error samples: Label them as "reasonably wrong" (humans would also struggle) vs "unreasonably wrong" (model/data problem).
  • Errors → Actions: Fix data (labeling errors), add data (blind spot classes), modify the model (mistaking background for landmarks, see Interpretability and Fairness).

5. A/B Testing and Online Evaluation ​

There's often a gap between offline metrics and online performance (distribution shift, latency, interaction effects). The gold standard before deployment is A/B testing:

StageQuestionMethod
OfflineAre model candidates good?Test set + confidence intervals
Shadow modeHow does the new model perform online?Run the new model alongside production, log only, no decisions
Small-traffic A/BDo business metrics improve?Route 5%–10%, run for 1–2 weeks
Full rolloutLong-term effectMonitor drift and have a rollback switch

A/B testing key points:

  • Predefine success metrics and minimum effect (e.g., conversion rate up 0.5% with a confidence interval excluding 0). Avoid "storytelling with data."
  • Watch out for Simpson's paradox: Overall no improvement but every subgroup improves (or vice versa) — always analyze by strata (new vs returning users, region, etc.).
  • Pay attention to the long tail: A model with good average metrics but degraded p99 shouldn't ship.
  • Have a rollback switch: If the new model breaks something, be able to revert with one click.

Continuous online evaluation involves monitoring, drift detection, and model update cadence — that falls under MLOps and Model Deployment.

6. Human Evaluation for Generative Models ​

Automated metrics for generative models (text, image, audio) like BLEU, FID, and Inception Score correlate poorly with human perception. You must supplement with structured human evaluation (basics in Generative Models and Diffusion Models and Generative AI):

  • Dimensional scoring: Don't rate "good or bad"; score by dimensions (relevance, fluency, factual accuracy, harmfulness, each on a 1–5 scale). Define scoring criteria, use multiple raters, and report inter-rater agreement (e.g., Cohen's kappa).
  • Paired preference: Ask raters to choose one (Model A vs Model B output). This is more stable than absolute scores and most closely resembles A/B testing.
  • Error taxonomy: Categorize generation errors (hallucination, repetition, instruction deviation), track proportions — this is more actionable than a single score.
  • Scale and cost: Human evaluation is expensive and slow; use reasonable sampling (~200–500 per config) plus LLM-based pre-screening (see Large Language Models (LLM)).

7. Case Study: Complete Image Classification Evaluation Pipeline ​

Bringing it all together into runnable code:

python
import numpy as np
import torch, torch.nn as nn
from torch.utils.data import DataLoader
from sklearn.metrics import (confusion_matrix, classification_report,
                             precision_recall_fscore_support)
import matplotlib.pyplot as plt

def full_evaluation(model, loader, device):
    """Returns (probs, preds, labels) and prints a full evaluation report"""
    model.eval()
    probs_list, preds_list, labels_list = [], [], []
    with torch.no_grad():
        for x, y in loader:
            logits = model(x.to(device))
            probs = torch.softmax(logits, dim=-1)
            probs_list.append(probs.cpu().numpy())
            preds_list.append(probs.argmax(-1).cpu().numpy())
            labels_list.append(y.numpy())
    probs = np.concatenate(probs_list)
    preds = np.concatenate(preds_list)
    labels = np.concatenate(labels_list)
    return probs, preds, labels

def evaluate_with_report(model, test_loader, classes, device, n_runs=5):
    """Multiple runs + confidence intervals + per-class metrics + confusion matrix + error analysis"""
    all_run_accs = []
    for seed in range(n_runs):
        torch.manual_seed(seed)
        # Note: same seed only re-runs randomness during evaluation; training randomness is handled in the experiment script
        probs, preds, labels = full_evaluation(model, test_loader, device)
        all_run_accs.append((preds == labels).mean())
        if seed == 0:
            print("=== per-class report ===")
            print(classification_report(labels, preds, target_names=classes, digits=4))
            cm = confusion_matrix(labels, preds)
            plt.imshow(cm, cmap="Blues")
            plt.colorbar()
            plt.title("Confusion matrix")
            plt.savefig("confusion_matrix.png")

    mean_acc = np.mean(all_run_accs)
    std_acc = np.std(all_run_accs, ddof=1)
    print(f"\naccuracy over {n_runs} runs: {mean_acc:.4f} ± {std_acc:.4f}")

    # Error analysis: high-confidence misclassifications ("confidently wrong")
    wrong = preds != labels
    conf_wrong = probs[wrong].max(axis=-1)
    top_idx = np.argsort(conf_wrong)[-10:]      # Top 10 highest-confidence errors
    print("top confident mistakes (index, true, pred, conf):")
    for i in top_idx:
        print(f"  {i}: true={labels[i]}, pred={preds[i]}, conf={conf_wrong[i]:.3f}")

    return mean_acc, std_acc, cm

Key usage notes:

  • Use the same test_loader definition for evaluation and training (same transforms, same data). Otherwise, you're not evaluating the model that was trained.
  • Average across multiple runs with mean ± std to make error bars like "±0.3%" visible.
  • Per-class report directly exposes weaknesses in imbalanced classes.
  • High-confidence error list is the seed for error analysis — manually inspecting these 10 samples often reveals data issues or model blind spots.

The artifacts from this pipeline (scores, confidence intervals, confusion matrix, error samples) form the evidence package for whether a model is "ready for deployment," and they're also the material for the "Results" section of Portfolio Projects.

Further Reading ​

References ​