Skip to content

Building a Model Evaluation Pipeline from Scratch

Quick overview Model evaluation isn't about computing a few metrics; it's a systematic method designed around business goals. Starting from evaluation system design, this article gives complete code and engineering practices for classification, regression, cross-validation, error analysis, class imbalance, and LLM evaluation.

Building a Model Evaluation Pipeline from Scratch ​

Let's start with an uncomfortable truth: no matter how beautiful a model's offline metrics (accuracy, AUC, F1) are, they can't prove it's useful online. Offline metrics are "proxies"; business results are the "real" — between them lie three full layers of uncertainty: data distribution, interaction feedback, and user behavior changes. This isn't negating offline evaluation; quite the opposite — precisely because there's a gap between proxy and reality, we need to treat evaluation itself as an engineering problem to take seriously.

This article doesn't cover "a catalog of metric formulas"; instead, it covers how to build an evaluation system from scratch: from evaluation system design, complete code for classification and regression, the right posture for cross-validation, error analysis, class imbalance, LLM application evaluation, and engineering implementation. After reading, you'll get an evaluation checklist you can copy directly into your own project.

I. Evaluation System Design: Clarify Four Things Before Writing Code ​

Almost every failed ML project can be traced to the same starting point: the wrong metric was chosen, or no metric was chosen at all. Many people start with accuracy_score, see 98%, and everyone's happy — until they discover the model missed every rare-disease patient, or sent all normal emails to spam. So before writing any evaluation code, clarify these four things first.

1. Define Business Goals First, Then Choose Metrics ​

The first principle of evaluation: business goals determine metrics, and metrics in turn constrain modeling. This four-level relationship can be drawn as a pyramid:

        Business goal (quantifiable, understood by boss and business stakeholders)
              ▲   mapped to
        Proxy metric (offline-computable, e.g., AUC, F1)
              ▲   constrains
        Evaluation protocol (how data is split, how experiments are repeated)
              ▲   constrains
        Statistical properties of each metric (variance, sensitivity to imbalance)

Top-down is "why," bottom-up is "how to guarantee." The same model, with different business goals, uses completely different metrics:

Business ScenarioQuantified Form of Business GoalCore Metric to WatchOne-Line Intuition
Spam filteringCost of misclassifying one normal email = losing a userPrecision (low FP)Better to miss spam than flag good mail
Disease screeningCost of missing one patient = delayed treatmentRecall (low FN)Better to over-check than miss
Credit risk controlBalance of default rate, approval rate, yieldAUC / KS + business stratified statsRanking ability determines lending strategy
Recommendation rankingCTR, conversion rate, exposure utilizationNDCG@k, MAPWhat's ranked higher matters more
Price predictionLarge deviations cause losses, small deviations are fineMAE or MAPE (not MSE)Don't let a few outliers hijack the model

The most common mistake here is forcing classification metrics (accuracy) onto ranking scenarios, or forcing regression metrics (MSE) onto businesses with large outlier amounts. When choosing metrics, ask yourself three questions:

  1. Is it monotonically related to the business goal? — does a 1% accuracy increase mean profit increases? Often, it doesn't.
  2. Is it sensitive to data distribution changes? — on imbalanced data, accuracy is drowned by the majority class (see Section VI).
  3. Is the variance high? — on small test sets, AUC's confidence interval can be so wide you question everything.

2. Define Baselines: Metrics Without Baselines Are Meaningless ​

"92% accuracy" — good or bad? Without a reference frame, there's no way to judge. Before investing in any complex model, hit at least three baselines:

Baseline TypeApproachPurpose
Majority class / random baselineAlways predict majority class, or uniform random predictionThe "difficulty floor" of the data itself
Heuristic baselineLast period's data, rule-based guessing, mean predictionDetermine how much increment ML actually brings
Lightweight model baselineLogistic regression, single decision treeDetermine whether complex model gains justify costs

Baselines aren't shameful; they're the measure zero point of the evaluation system. A rule of thumb: if a complex model is only 2 points above majority-class baseline, first suspect the evaluation pipeline, then suspect the model.

3. Define Comparison Protocols: Make All Experiments Reproducible and Comparable ​

Evaluation isn't one run and done; it's a multi-round process. Without a protocol, round two experiments can't be compared to round one. A minimum viable protocol contains five items:

  • Lock one test set from day one; no one can change it for "better looks";
  • Fix random seeds: fix model, split, sampling — ensure differences come from methods, not luck;
  • Repeat experiments multiple times, report mean ± std, not just single results;
  • Record data version and code version — the only answer to "did this improvement come from changing data or changing model?";
  • Commit evaluation scripts alongside training scripts; evaluation code must never be separated from the model.

Test set contamination is evaluation's #1 accident

Any behavior that involves "looking at the test set, tuning on the test set, reverse-engineering features from the test set, early-stopping on the test set" creates a fake model — strong on paper, crashes on deployment. The professional approach: never touch the test set, iterate on the validation set only, then use the test set for one-time final confirmation. This discipline is repeatedly emphasized in Model Evaluation and Validation.

II. Classification Evaluation: From Confusion Matrix to Threshold Selection ​

Let's first create labeled data; all classification code here can be run directly. We use a 9:1 imbalanced binary classification dataset — not artificially balanced, because that's exactly what Section VI addresses.

python
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import (
    confusion_matrix, classification_report,
    roc_curve, roc_auc_score,
    precision_recall_curve, average_precision_score,
    f1_score, precision_score, recall_score,
)

# Generate 5000 samples, 20 features, positive class at 10%
X, y = make_classification(
    n_samples=5000, n_features=20, n_informative=12,
    n_redundant=4, weights=[0.9, 0.1], random_state=42,
)

# Stratified split: ensure positive ratio is consistent in train/test
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, stratify=y, random_state=42,
)

model = RandomForestClassifier(n_estimators=200, random_state=42)
model.fit(X_train, y_train)

y_prob = model.predict_proba(X_test)[:, 1]  # predicted probability for positive class
y_pred = model.predict(X_test)              # hard classification at default threshold 0.5

1. The Confusion Matrix Is the Source of All Classification Metrics ​

All information is hidden in this four-cell table:

                  Predicted Positive(Pred+)    Predicted Negative(Pred-)
Actual Positive(Actual+)   TP True Positive         FN False Negative (miss)
Actual Negative(Actual-)   FP False Positive (alarm) TN True Negative
python
cm = confusion_matrix(y_test, y_pred)
print("Confusion matrix:\n", cm)

# TN  FP
# FN  TP
tn, fp, fn, tp = cm.ravel()
print(f"TP={tp}  FN={fn}  FP={fp}  TN={tn}")

Looking at cm alone reveals problems accuracy can't show: if the positive class is only 10%, predicting all negative gets 90% accuracy, but TP=0. So the first thing is always to print the confusion matrix, not just print one total score.

2. Six Basic Metrics, One Table to Clarify ​

MetricFormulaIntuitionPrimary Use Case
Accuracy(TP+TN)/(TP+FN+FP+TN)Proportion predicted correctly overallWhen classes are balanced
PrecisionTP/(TP+FP)Of predicted positives, how many are rightWhen false alarms are costly
RecallTP/(TP+FN)Of actual positives, how many were caughtWhen misses are costly
F12·P·R/(P+R)Harmonic mean of P and RWhen both matter
Fβ(1+β²)·P·R/(β²·P+R)Give R β× weightWhen favoring one side
AUC / APSee belowRanking ability / average precision for positivesThreshold-independent overall comparison
python
print(classification_report(y_test, y_pred))
print(f"F1 = {f1_score(y_test, y_pred):.3f}")

classification_report gives P/R/F1 and per-class sample counts in one go; it's the main tool for daily iteration.

3. Threshold Is Part of Evaluation, Not a Model Parameter ​

Most classifiers output probabilities, not labels; predict() just cuts at 0.5. When business costs are asymmetric (e.g., missing a diagnosis is far worse than a false alarm), 0.5 is often not the optimal cut point. The right posture: treat the threshold as a hyperparameter to scan:

python
for threshold in [0.3, 0.4, 0.5, 0.6, 0.7]:
    y_t = (y_prob >= threshold).astype(int)
    print(f"Threshold {threshold}:  Precision={precision_score(y_test, y_t):.3f}  "
          f"Recall={recall_score(y_test, y_t):.3f}  "
          f"F1={f1_score(y_test, y_t):.3f}")

Which threshold to choose depends on the business cost table: one missed diagnosis = how many yuan, one false alarm = how many yuan; find the cut point that minimizes total cost. This is the key step of "moving evaluation from inside the model to the business layer" — evaluation output should be a "threshold-metrics-cost" table, not a label array.

4. PR and ROC Curves: Plotting the Model's Full Spectrum ​

ROC looks at the trade-off between true positive rate and false positive rate at all thresholds; the PR curve looks at the trade-off between precision and recall at all thresholds. AUC is the area under the curve, measuring "ranking ability": the probability that, given a random positive and a random negative, the model gives the positive a higher score.

python
# ROC curve and AUC
fpr, tpr, roc_thresholds = roc_curve(y_test, y_prob)
roc_auc = roc_auc_score(y_test, y_prob)

# PR curve and AP (Average Precision, area under PR curve)
precision, recall, pr_thresholds = precision_recall_curve(y_test, y_prob)
ap = average_precision_score(y_test, y_prob)

print(f"ROC-AUC = {roc_auc:.3f}   PR-AUC(AP) = {ap:.3f}")

fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 4))

ax1.plot(fpr, tpr, label=f"AUC={roc_auc:.3f}")
ax1.plot([0, 1], [0, 1], "k--", label="Random baseline")
ax1.set(xlabel="FPR", ylabel="TPR", title="ROC Curve")
ax1.legend()

ax2.plot(recall, precision, label=f"AP={ap:.3f}")
ax2.axhline(y_train.mean(), color="k", ls="--",
            label="Positive ratio baseline")
ax2.set(xlabel="Recall", ylabel="Precision",
        title="PR Curve")
ax2.legend()
plt.tight_layout()
plt.show()

Note that ROC's random baseline is the y=x diagonal line (independent of positive ratio), while PR curve's baseline is the horizontal line at the positive ratio. When positives are only 10%, PR curve's baseline is 0.1 — this is exactly the key to "ROC lies" in Section VI.

5. Package Everything Above Into One Function ​

Daily iteration needs "one command output full evaluation." Consolidating scattered code into a reusable evaluator is the first step of evaluation engineering:

python
def evaluate_binary_classifier(y_true, y_prob, name="Model"):
    """Output a full offline evaluation report for a binary classifier."""
    results = {}
    results["ROC-AUC"] = roc_auc_score(y_true, y_prob)
    results["AP"] = average_precision_score(y_true, y_prob)

    # Scan thresholds, find P/R for best F1
    p, r, th = precision_recall_curve(y_true, y_prob)
    f1s = 2 * p * r / (p + r + 1e-9)
    best = np.argmax(f1s)
    results["best_threshold"] = th[best]
    results["best_P"] = p[best]
    results["best_R"] = r[best]
    results["best_F1"] = f1s[best]
    results["confusion_matrix"] = confusion_matrix(
        y_true, (y_prob >= results["best_threshold"]).astype(int))

    print(f"===== {name} =====")
    for k, v in results.items():
        if k != "confusion_matrix":
            print(f"{k}: {v:.3f}")
    print("Confusion matrix (at best F1 threshold):\n", results["confusion_matrix"])
    return results

A clear evaluation function should be output information volume > input code volume, and always hangable into regression tests.

III. Regression Evaluation: More Than Just MSE ​

Regression evaluation seems simple on the surface (one error number), but a single error number is exactly the biggest pitfall — MSE is hijacked by outliers, R² is deceived by variance. The approach: metrics as foundation + residual diagnosis.

1. Three Core Metrics + One Yes/No Question ​

MetricFormulaPropertyWhen It Dominates
MSE(1/n)Σ(yᵢ-ŷᵢ)²Penalizes large errors (square amplifies)Large errors unacceptable
RMSE√MSESame units as y, intuitiveNeed to compare directly with y
MAE(1/n)Σ|yᵢ-ŷᵢ|Robust to outliersError distribution has heavy tails
R²1 − SS_res/SS_totExplained variance relative to mean baselineRelative comparison between models

A key yes/no question: when outliers come from "normal but extreme business" (large orders, extreme weather), use MAE; when outliers represent "system failures / dirty data," clean first rather than switching metrics.

python
from sklearn.datasets import make_regression
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score

X, y = make_regression(n_samples=3000, n_features=10, noise=15, random_state=42)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=42)

reg = RandomForestRegressor(n_estimators=200, random_state=42)
reg.fit(X_tr, y_tr)
y_hat = reg.predict(X_te)

mse  = mean_squared_error(y_te, y_hat)
rmse = np.sqrt(mse)
mae  = mean_absolute_error(y_te, y_hat)
r2   = r2_score(y_te, y_hat)

print(f"MSE={mse:.1f}  RMSE={rmse:.1f}  MAE={mae:.1f}  R²={r2:.3f}")

If RMSE is clearly larger than MAE (e.g., 2×), the error distribution's right tail is heavy — a few samples contribute most of the error, and this is exactly where the next step, residual analysis, digs.

2. Residual Analysis: Evaluation's X-Ray ​

Residuals = true value − predicted value. A good model's residuals should have mean 0, stable variance, and no relationship with features or predictions. Four plots are enough:

python
residuals = y_te - y_hat

fig, axes = plt.subplots(2, 2, figsize=(11, 8))

# (1) Residuals vs predicted: should be a random band, no trumpet / bend shape
axes[0, 0].scatter(y_hat, residuals, s=8, alpha=0.5)
axes[0, 0].axhline(0, color="r", lw=1)
axes[0, 0].set(xlabel="Predicted", ylabel="Residual", title="Residuals vs Predicted")

# (2) Residual histogram: should be approximately normal, centered at 0
axes[0, 1].hist(residuals, bins=50, alpha=0.7)
axes[0, 1].axvline(0, color="r", lw=1)
axes[0, 1].set(xlabel="Residual", ylabel="Count", title="Residual Distribution")

# (3) Residuals vs most important feature: check for missed non-linear relationships
imp = np.argsort(-reg.feature_importances_)[0]
axes[1, 0].scatter(X_te[:, imp], residuals, s=8, alpha=0.5)
axes[1, 0].axhline(0, color="r", lw=1)
axes[1, 0].set(xlabel=f"Most important feature #{imp}", ylabel="Residual",
               title="Residuals vs Most Important Feature")

# (4) Cumulative residual ratio: see what proportion of samples contribute most error
axes[1, 1].plot(np.sort(np.abs(residuals))[::-1].cumsum()
                / np.abs(residuals).sum())
axes[1, 1].set(xlabel="Sample index (sorted by error descending)", ylabel="Cumulative error ratio",
               title="Error Concentration Curve")
plt.tight_layout()
plt.show()

Four typical symptoms and corresponding prescriptions:

Residual SymptomMeaningAction
Trumpet shape (residuals increase with predicted value)Heteroscedasticity, possibly missing log transformModel y after log
Bent / opening curveMissed non-linear terms or feature interactionsAdd feature interactions, switch to tree models
Residual mean significantly deviates from 0Systematic bias (underfitting or features leak directionally)Check features, increase capacity
Error concentrates in few samplesHeavy tails, a few "large orders" dominate errorFall back on MAE for evaluation, chase these 5% separately

High R² but systematic patterns in residuals means the model is "accurate on the surface, biased inside." The essence of regression evaluation isn't looking at one score; it's making the misalignment between model and data visible. How bias-variance and regularization further affect errors, see Overfitting and Regularization.

IV. The Right Way to Use Cross-Validation ​

A single train/test split result has randomness — change the seed, and AUC might jump from 0.83 to 0.87. Cross-validation (CV) uses the mean ± std of multiple splits to give a stable estimate, while using every sample.

1. StratifiedKFold: The Standard Choice for Classification ​

python
from sklearn.model_selection import StratifiedKFold, cross_val_score

skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

scores = cross_val_score(model, X, y, cv=skf, scoring="roc_auc")
print(f"5-fold ROC-AUC: {scores.mean():.3f} ± {scores.std():.3f}")

Two key points:

  • shuffle=True + random_state=42: without shuffling, if data is ordered by class, each fold's class distribution will be severely imbalanced;
  • Stratified rather than plain KFold: ensures each fold's positive/negative ratio matches the full data, especially critical on small, imbalanced datasets.

cross_val_score is for quick estimation, but to get per-fold prediction probabilities for error analysis, loop manually:

python
oof = np.zeros(len(X))  # out-of-fold prediction probabilities
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for fold, (tr_idx, va_idx) in enumerate(skf.split(X, y)):
    clf = RandomForestClassifier(n_estimators=200, random_state=42)
    clf.fit(X[tr_idx], y[tr_idx])
    oof[va_idx] = clf.predict_proba(X[va_idx])[:, 1]
print(f"OOF AUC = {roc_auc_score(y, oof):.3f}")

"Out-of-fold predictions" (OOF) is the golden data source for Section V error analysis: each sample's prediction comes from a model that never saw it, so it's safe to use for error pattern statistics.

2. Leakage: CV's Most Concealed Killer ​

CV's statistical validity rests on one premise: each fold's training data contains only information from that fold's training set. The three most common leakages:

Leakage SourceExampleConsequence
Preprocessing done outside CVStandardize / mean-impute on full data first, then do CVValidation set information leaks into training, inflated metrics
Feature selection leaksSelect features on full data before entering CVSelected features that "leak the future"
Random split for time dataUse day 3 data to train, day 1 data to validateModel "sees the future," time-series scenario fails

The standard solution is Pipeline — let each preprocessing step be inside CV, re-fit per fold:

python
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

pipe = Pipeline([
    ("scaler", StandardScaler()),          # re-fit inside each fold
    ("clf", LogisticRegression(max_iter=2000)),
])
scores = cross_val_score(pipe, X, y, cv=skf, scoring="roc_auc")
print(f"Pipeline 5-fold AUC: {scores.mean():.3f} ± {scores.std():.3f}")

3. Group and Time-Series: When Samples Aren't Independent ​

If your data has group dependency structure (multiple behaviors from the same user, multiple orders from the same batch of goods, multiple patches from the same image), plain KFold splits the same group's data across train and validation — the model has "seen the same user," so evaluation is naturally inflated. Use GroupKFold to guarantee the same group's samples always stay in the same fold:

python
from sklearn.model_selection import GroupKFold

# group array: which user/batch each sample belongs to
group = np.arange(len(X)) // 10   # example: every 10 samples from the same user
gkf = GroupKFold(n_splits=5)
scores = cross_val_score(model, X, y, cv=gkf, groups=group, scoring="roc_auc")
print(f"GroupKFold AUC: {scores.mean():.3f} ± {scores.std():.3f}")

Time-series data uses TimeSeriesSplit (only "past to train, future to validate"), and note two rules: training always predates validation; validation must not contain future information from the training set (e.g., when features include "historical mean," compute them point-by-point in a rolling manner, not all at once).

4. Nested CV: When CV Is Used for Both Tuning and Evaluation ​

If you use CV for both tuning (GridSearchCV) and reporting final metrics, the second set of scores is contaminated by the tuning process (the validation set has been seen too many times). The rigorous approach is nested CV: outer CV estimates generalization, inner CV handles tuning.

python
from sklearn.model_selection import GridSearchCV

param_grid = {"n_estimators": [50, 200], "max_depth": [5, None]}
inner = GridSearchCV(RandomForestClassifier(random_state=42),
                     param_grid, cv=StratifiedKFold(3, shuffle=True, random_state=1))
outer_scores = cross_val_score(inner, X, y,
                               cv=StratifiedKFold(5, shuffle=True, random_state=2),
                               scoring="roc_auc")
print(f"Nested CV AUC: {outer_scores.mean():.3f} ± {outer_scores.std():.3f}")

Nested CV doubles computation; you don't need it every daily iteration, but at key decision points ("what is this model truly worth"), running it once is worth it. The trade-off between CV and bootstrap for small samples, refer to Kohavi's classic study (see references).

V. Error Analysis and the Iteration Loop ​

Evaluation isn't the end; it's the input for the next iteration. Without error analysis, evaluation is just "reporting a score"; with it, evaluation becomes "finding where the model improves next."

1. What Error Analysis Is ​

In one sentence: on the validation set (or OOF predictions), find the samples the model got wrong, understand why, and based on this decide the next action. The core tool is "turn evaluation results into a queryable table."

python
# Turn the test set into an analyzable DataFrame
df_results = pd.DataFrame({
    "y_true": y_test,
    "y_prob": y_prob,
    "y_pred": y_pred,
    "is_error": y_test != y_pred,
    "abs_error": np.abs(y_test - y_pred),          # for regression
    "confidence": np.where(
        y_prob >= 0.5, y_prob, 1 - y_prob),         # model confidence
})

2. Group Statistics by Error Type ​

With the table, error analysis is a series of groupby operations. Below is a directly usable grouping matrix:

python
# ① By business field: which channel/time period/customer segment has the highest error rate?
df_results["group"] = X_test[:, 0] > 0   # example: split by a feature into two groups
group_stats = (df_results.groupby("group")["is_error"]
               .agg(errors="sum", count="count", error_rate="mean"))
print(group_stats)

# ② By confidence tier: what proportion of errors are "model was very confident and wrong"?
df_results["conf_bin"] = pd.cut(df_results["confidence"],
                                bins=[0, 0.6, 0.8, 0.9, 1.0])
print(df_results.groupby("conf_bin", observed=True)["is_error"]
      .agg(error_rate="mean", count="count"))

# ③ Average confidence of error samples vs correct samples
print(df_results.groupby("is_error")["confidence"].mean())

Three typical findings and their directed actions:

Analysis FindingImplied Root CauseAction
One segment / channel has error rate far above averageThis subset lacks features or samplesAdd features, add data for this subset
Error samples all have high confidenceModel overconfident / poorly calibratedUse CalibratedClassifierCV for probability calibration
Error samples happen to be "rare sub-classes"Class imbalance / data sparsitySee Section VI's resampling and weighting approaches
Error samples share a common time periodTime-correlated driftAdd time features, switch to time-based split

The right way to do error analysis

Don't just look at summary numbers; personally flip through 20–30 raw error samples yourself. Numbers tell you "where errors are concentrated," samples tell you "why the errors are absurd" — e.g., you thought the model learned semantics, but it's actually copying repeated sentences from the training set. This cognitive gap only appears when you see raw samples. Interpretability tools (SHAP, partial dependence plots) can turn this "manual sample-flipping" into systematic attribution; see Interpretability and Fairness.

3. Turning Analysis Into Action: The Evaluation–Iteration Loop ​

Error analysis produces hypotheses, not conclusions. Each analysis converges on a "finding → hypothesis → action" table; only do the one with the biggest impact, then re-evaluate:

Evaluate (fixed protocol) → Error analysis (group stats + flip samples)
        ↑                                        ↓
  Re-evaluate, compare to previous round     Form hypothesis
        ↑                                        ↓
  Record: changed data / features / model / threshold  →  Execute one minimal change

A practical iteration discipline:

  • Change one variable at a time: add features, switch models, adjust thresholds — do them separately, or improvement attribution is unclear;
  • Every iteration has a comparison record: what changed, before/after val metrics, whether new errors were introduced;
  • Calibrate changes back to "business goals": val AUC went up 0.01, but the online conversion proxy metric didn't move → the improvement is ineffective;
  • Acknowledge "evaluation saturation": if two to three rounds produce no significant gains, stop and do a big review, rather than continuing to fine-tune.

The full skeleton of this loop can be mapped onto the workflow in Building an ML Project from Scratch; the most common mistakes in iteration (repeatedly tuning on validation set, metrics decoupled from business, etc.) have a ready-made avoidance list in Common Pitfalls and Anti-Patterns.

VI. Class Imbalance Evaluation: Why PR-AUC Usually Beats ROC ​

Imbalance isn't "data has a problem"; it's the default state of most real businesses: fraud, disease, clicks, failures — rare events always occupy the minority. Under imbalance, the choice of evaluation metric directly determines whether you see the truth or an illusion.

1. Why ROC Can "Lie" ​

ROC's horizontal and vertical axes are TPR and FPR, both of which are ratios with their respective classes as denominators — it treats positive and negative classes symmetrically, so ROC shape changes very little whether positives are 1% or 50%. But PR curve's vertical axis is precision, whose baseline is the positive ratio: when positives are 1%, random guessing gives 0.5 ROC-AUC, but PR baseline is only 0.01.

Consider an extreme but real example (positives at 1%, 10,000 test samples):

ScenarioROC-AUCAP (PR-AUC)
Model predicts all samples negative0.5 (equal to random)Undefined (no positive predictions)
Model catches 60% positives, 50 false alarmsCan reach 0.98+~0.55
Model catches 10% positives, 3 false alarmsCan also reach 0.95~0.25

In the same experiment, ROC-AUC looks pleasantly high, but AP is shockingly low — because ROC is "diluted" by the huge negative class, while the PR curve directly reflects "when you say it's positive, how likely are you right?" The lower the positive ratio, the more ROC and PR conclusions diverge. For rigor, recommend looking at both, but when the business cares about "are positives found correctly," take PR-AUC/AP as the standard. Theoretical details are in Davis & Goadrich and Saito & Rehmsmeier's papers (references).

2. The Correct Posture for Imbalanced Evaluation ​

  • Report PR-AUC/AP, simultaneously print confusion matrix absolute values (absolute TP/FN/FP out of 10k are more business-meaningful than ratios);
  • Use Fβ instead of F1: increase β when misses are costly;
  • Report P/R pairs at business thresholds: don't chase "best F1"; instead select threshold from a cost table, then report P/R;
  • Don't just look at accuracy — it's dominated by the 99% majority class.

3. Mitigation Approaches, In Order ​

Fix evaluation first, then data, then model only last. Get the order wrong, and everything upstream is distorted:

python
# ① Don't change data, first weight the loss function (internal model solution)
from sklearn.ensemble import RandomForestClassifier
weighted = RandomForestClassifier(n_estimators=200,
                                  class_weight="balanced", random_state=42)
weighted.fit(X_train, y_train)
print("class_weight=balanced AP:",
      average_precision_score(y_test, weighted.predict_proba(X_test)[:, 1]))

# ② Oversample minority class (resampling approach, must be done inside CV)
from imblearn.over_sampling import SMOTE
from imblearn.pipeline import Pipeline as ImbPipeline

smote_pipe = ImbPipeline([
    ("smote", SMOTE(random_state=42)),
    ("clf", RandomForestClassifier(n_estimators=200, random_state=42)),
])
smote_scores = cross_val_score(smote_pipe, X, y, cv=skf, scoring="average_precision")
print(f"SMOTE 5-fold AP: {smote_scores.mean():.3f} ± {smote_scores.std():.3f}")

Resampling must never appear in the validation pipeline

SMOTE or undersampling can only happen inside training folds — if you SMOTE on full data first, then split, synthesized minority samples appear in both train and validation, and validation metrics are inflated by their own "duplicates." The code above uses imblearn.pipeline.Pipeline specifically to guarantee SMOTE executes per-fold inside CV.

Beyond resampling, there's collecting more positive samples (always most effective), changing the problem to anomaly detection / ranking (when positives are single digits), and cost-sensitive learning (writing false alarm/miss costs directly into the objective function). Remember: data itself being imbalanced isn't the disease; the evaluation approach being wrong is the disease — class_weight and threshold selection are often more effective than fancy resampling.

VII. LLM Application Evaluation: A Brief Intro ​

LLM applications (RAG, Agents, summarization, code generation) evaluation differs fundamentally from classical ML: outputs are open-ended text, correctness is highly subjective, and errors aren't "right/wrong" but "which part is wrong." But the four steps of evaluation system design are fully applicable — just the implementation differs. This section is a quick survey; full solutions are in Large Language Model Application Evaluation.

1. Constructing the Eval Set ​

LLM evaluation also needs "train / dev / test" layering, but what's split here isn't model weights, but prompts and system design:

SetPurposeRecommended Size
Dev setIterate prompts, tune RAG parameters50–200 samples
Test setFinal acceptance, ideally close to real distribution200–1000 samples
Golden setAuto-regression, fixed unchanged100–500 samples

Collection priority: online real requests > manually constructed typical scenarios > mined failure samples from user feedback. A common mistake is the eval set only containing "smooth questions," so the model can't even handle "a user asked an irrelevant question" — the eval set must deliberately include edge cases (ambiguity, missing context, long-tail dialects, malicious input).

2. Human Evaluation: The Only Trustworthy Scale ​

Automated metrics still can't replace humans in judging "is this summary accurate, is this response offensive?" To land a reproducible human evaluation:

  • Define a scorecard: correctness, faithfulness (faithfulness to context), usefulness, style/tone, each 1–5 points with clear anchors;
  • Multiple people score, compute inter-annotator agreement; re-discuss rules for samples with disagreement;
  • Re-score the same batch after each change, avoiding "different batches, different standards" drift.

Human evaluation is slow and expensive, so its position is as the calibration ruler for automated metrics, not the main tool for daily iteration.

3. Automated Metrics: Can Save Manpower, But Don't Worship ​

TypeExamplesPurposeLimitation
Text similarityBLEU, ROUGESummarization, translationCompletely ineffective for "semantically correct but differently worded"
Component-level metricsRAG retrieval Recall@k, hit rateEvaluate RAG retrieval layer in isolationRequires layered evaluation design
LLM-as-judgeUse a strong model (e.g., GPT-4) to score a weaker modelFaithfulness, relevanceSelf-preference bias

Three landing key points for LLM-as-judge: provide structured scoring criteria (far more stable than free-form comments); run judge consistency checks (score a batch of samples with both humans and judges; if correlation is below threshold, it's unreliable); prioritize scoring dimensions with anchors like "factuality / relevance" over "writing quality." More nuanced issues on hallucination and eval bias are in the Large Language Models deep-dive.

4. Layered Evaluation: RAG Applications Should Be Split Into at Least Three Layers ​

For an end-to-end RAG Q&A system, when the overall metric is poor, you have no idea whether retrieval didn't retrieve or the model didn't answer correctly. Evaluate each layer separately:

User question ──► Retrieval layer (Recall@k: was the correct document recalled?)
                │
                ▼
           Generation layer (Faithfulness: does the answer faithful to the retrieved docs?)
                │
                ▼
           Interaction layer (Usefulness: did the answer solve the user's problem?)

Each layer uses different metrics, different failure modes, and different optimization actions. The value of a single end-to-end score is far less than this three-layer health check table.

VIII. Evaluation Engineering: Making Evaluation a Systematic Part ​

Doing evaluation once isn't hard; the challenge is automating and reproducing evaluation after every change, and continuously monitoring post-deployment. This belongs to the realm of MLOps and Production Deployment; here are three minimum must-do actions.

1. Evaluation Scripts In Repo, Same Repo as Model Code ​

Evaluation isn't a casual brush in Jupyter; it's versioned code. Recommended directory structure:

project/
├── model/            # training code
├── evals/
│   ├── metrics.py    # metrics and evaluation functions (e.g., evaluate_binary_classifier from Section II)
│   ├── datasets/     # test sets, golden sets (with version numbers)
│   ├── run_eval.py   # one-click eval entry, outputs JSON/Markdown reports
│   └── tests/        # unit tests for the evaluation itself
└── reports/          # historical evaluation reports, traceable

run_eval.py outputs a report with metrics, data version, code commit, random seed, run time every time — this is the only evidence chain for "what changed last round, and did it improve?"

2. Use Golden Sets for Regression Testing ​

Turn evaluation from "run when someone remembers" to "run automatically after every change":

  • Pick 100–500 golden samples; after evaluation, assert key metrics don't fall below threshold (e.g., assert ap >= 0.40);
  • Integrate into CI: changing model code, features, or data scripts all trigger evaluation; metrics below threshold → build fails;
  • Make "never regress" an engineering convention: models before deployment must pass all golden assertions, and not fall below the current online model's metrics on the test set.
python
# Skeleton of run_eval.py (excerpt)
def check_regression(results, thresholds):
    failures = []
    for metric, limit in thresholds.items():
        if results[metric] < limit:
            failures.append(f"{metric}={results[metric]:.3f} < {limit}")
    if failures:
        raise RuntimeError("Regression test failed: " + "; ".join(failures))

3. Post-Deployment Monitoring: From "One-Time" to "Continuous" ​

The premise of offline evaluation is "test set represents the future," but distributions drift online. Deploy at least three monitoring channels:

Monitoring ItemMethodWarning Signal
Data driftFeature distribution (PSI / KL divergence) daily comparisonPSI > 0.2 needs review
Label driftOnline label / result distribution changesSudden shift in positive rate
Business proxy metricsOnline conversion rate, success rate, user complaintsMetrics diverge from offline expectations

Complete practices for drift detection and model retraining cycles are in MLOps and Production Deployment. One principle running throughout: online monitoring metrics must be the same as to offline evaluation metrics — otherwise you'll simultaneously have two mutually contradictory truths.

IX. Trade-offs and Considerations ​

There's no free lunch in evaluation systems; every bit of rigor has a cost. These trade-offs should be made proactively and documented, not passively borne.

  • Metric fidelity vs. iteration speed: nested CV, repeated experiments, thousand-person scoring are the most rigorous, but each iteration takes hours or days. Use lightweight evaluation (single StratifiedKFold + proxy metrics) for daily iteration; use heavy evaluation (nested CV, human scoring) at key decision points (deployment, architecture changes). Two-tier evaluation is a common solution balancing both.
  • Test set size vs. estimation precision: larger test sets = more stable metrics, but less data for training. When data is limited, prefer to train set slightly smaller to keep the test set representative — the test set is the "final judge," and judges can't starve.
  • Metric granularity vs. interpretability: NDCG, KS, calibration error are more precise, but business stakeholders don't understand them; accuracy is universally understood but often misleading. Put both "coarse metrics the business understands" and "precise metrics engineers use" in the report, don't pick one.
  • Automation vs. human review: automated metrics are cheap but may be distorted (especially for LLM evaluation); human evaluation is credible but expensive. The compromise is automated screening + human spot-checking: automation runs the full set, sample by layer ("predicted failure / low confidence / edge cases") for human review.
  • The temptation to use validation as test set: after 50 iterations, the "validation set" has actually been seen by you 50 times. A honest evaluation system reserves an untouched final test set, and does one-time final confirmation after validation metrics saturate. See Common Traps in Evaluation.

A final overarching principle: the value of an evaluation system doesn't depend on how advanced its metrics are, but on whether it can stably answer "will my business improve after this model goes live?" Metrics are proxies; business is reality; evaluation's purpose is to make the gap between the two as small as possible, and known.

X. Further Reading ​

References ​