Theme
Building a Model Evaluation Pipeline from Scratch
Let's start with an uncomfortable truth: no matter how beautiful a model's offline metrics (accuracy, AUC, F1) are, they can't prove it's useful online. Offline metrics are "proxies"; business results are the "real" — between them lie three full layers of uncertainty: data distribution, interaction feedback, and user behavior changes. This isn't negating offline evaluation; quite the opposite — precisely because there's a gap between proxy and reality, we need to treat evaluation itself as an engineering problem to take seriously.
This article doesn't cover "a catalog of metric formulas"; instead, it covers how to build an evaluation system from scratch: from evaluation system design, complete code for classification and regression, the right posture for cross-validation, error analysis, class imbalance, LLM application evaluation, and engineering implementation. After reading, you'll get an evaluation checklist you can copy directly into your own project.
I. Evaluation System Design: Clarify Four Things Before Writing Code
Almost every failed ML project can be traced to the same starting point: the wrong metric was chosen, or no metric was chosen at all. Many people start with accuracy_score, see 98%, and everyone's happy — until they discover the model missed every rare-disease patient, or sent all normal emails to spam. So before writing any evaluation code, clarify these four things first.
1. Define Business Goals First, Then Choose Metrics
The first principle of evaluation: business goals determine metrics, and metrics in turn constrain modeling. This four-level relationship can be drawn as a pyramid:
Business goal (quantifiable, understood by boss and business stakeholders)
▲ mapped to
Proxy metric (offline-computable, e.g., AUC, F1)
▲ constrains
Evaluation protocol (how data is split, how experiments are repeated)
▲ constrains
Statistical properties of each metric (variance, sensitivity to imbalance)Top-down is "why," bottom-up is "how to guarantee." The same model, with different business goals, uses completely different metrics:
| Business Scenario | Quantified Form of Business Goal | Core Metric to Watch | One-Line Intuition |
|---|---|---|---|
| Spam filtering | Cost of misclassifying one normal email = losing a user | Precision (low FP) | Better to miss spam than flag good mail |
| Disease screening | Cost of missing one patient = delayed treatment | Recall (low FN) | Better to over-check than miss |
| Credit risk control | Balance of default rate, approval rate, yield | AUC / KS + business stratified stats | Ranking ability determines lending strategy |
| Recommendation ranking | CTR, conversion rate, exposure utilization | NDCG@k, MAP | What's ranked higher matters more |
| Price prediction | Large deviations cause losses, small deviations are fine | MAE or MAPE (not MSE) | Don't let a few outliers hijack the model |
The most common mistake here is forcing classification metrics (accuracy) onto ranking scenarios, or forcing regression metrics (MSE) onto businesses with large outlier amounts. When choosing metrics, ask yourself three questions:
- Is it monotonically related to the business goal? — does a 1% accuracy increase mean profit increases? Often, it doesn't.
- Is it sensitive to data distribution changes? — on imbalanced data, accuracy is drowned by the majority class (see Section VI).
- Is the variance high? — on small test sets, AUC's confidence interval can be so wide you question everything.
2. Define Baselines: Metrics Without Baselines Are Meaningless
"92% accuracy" — good or bad? Without a reference frame, there's no way to judge. Before investing in any complex model, hit at least three baselines:
| Baseline Type | Approach | Purpose |
|---|---|---|
| Majority class / random baseline | Always predict majority class, or uniform random prediction | The "difficulty floor" of the data itself |
| Heuristic baseline | Last period's data, rule-based guessing, mean prediction | Determine how much increment ML actually brings |
| Lightweight model baseline | Logistic regression, single decision tree | Determine whether complex model gains justify costs |
Baselines aren't shameful; they're the measure zero point of the evaluation system. A rule of thumb: if a complex model is only 2 points above majority-class baseline, first suspect the evaluation pipeline, then suspect the model.
3. Define Comparison Protocols: Make All Experiments Reproducible and Comparable
Evaluation isn't one run and done; it's a multi-round process. Without a protocol, round two experiments can't be compared to round one. A minimum viable protocol contains five items:
- Lock one test set from day one; no one can change it for "better looks";
- Fix random seeds: fix model, split, sampling — ensure differences come from methods, not luck;
- Repeat experiments multiple times, report mean ± std, not just single results;
- Record data version and code version — the only answer to "did this improvement come from changing data or changing model?";
- Commit evaluation scripts alongside training scripts; evaluation code must never be separated from the model.
Test set contamination is evaluation's #1 accident
Any behavior that involves "looking at the test set, tuning on the test set, reverse-engineering features from the test set, early-stopping on the test set" creates a fake model — strong on paper, crashes on deployment. The professional approach: never touch the test set, iterate on the validation set only, then use the test set for one-time final confirmation. This discipline is repeatedly emphasized in Model Evaluation and Validation.
II. Classification Evaluation: From Confusion Matrix to Threshold Selection
Let's first create labeled data; all classification code here can be run directly. We use a 9:1 imbalanced binary classification dataset — not artificially balanced, because that's exactly what Section VI addresses.
python
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import (
confusion_matrix, classification_report,
roc_curve, roc_auc_score,
precision_recall_curve, average_precision_score,
f1_score, precision_score, recall_score,
)
# Generate 5000 samples, 20 features, positive class at 10%
X, y = make_classification(
n_samples=5000, n_features=20, n_informative=12,
n_redundant=4, weights=[0.9, 0.1], random_state=42,
)
# Stratified split: ensure positive ratio is consistent in train/test
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, stratify=y, random_state=42,
)
model = RandomForestClassifier(n_estimators=200, random_state=42)
model.fit(X_train, y_train)
y_prob = model.predict_proba(X_test)[:, 1] # predicted probability for positive class
y_pred = model.predict(X_test) # hard classification at default threshold 0.51. The Confusion Matrix Is the Source of All Classification Metrics
All information is hidden in this four-cell table:
Predicted Positive(Pred+) Predicted Negative(Pred-)
Actual Positive(Actual+) TP True Positive FN False Negative (miss)
Actual Negative(Actual-) FP False Positive (alarm) TN True Negativepython
cm = confusion_matrix(y_test, y_pred)
print("Confusion matrix:\n", cm)
# TN FP
# FN TP
tn, fp, fn, tp = cm.ravel()
print(f"TP={tp} FN={fn} FP={fp} TN={tn}")Looking at cm alone reveals problems accuracy can't show: if the positive class is only 10%, predicting all negative gets 90% accuracy, but TP=0. So the first thing is always to print the confusion matrix, not just print one total score.
2. Six Basic Metrics, One Table to Clarify
| Metric | Formula | Intuition | Primary Use Case |
|---|---|---|---|
| Accuracy | (TP+TN)/(TP+FN+FP+TN) | Proportion predicted correctly overall | When classes are balanced |
| Precision | TP/(TP+FP) | Of predicted positives, how many are right | When false alarms are costly |
| Recall | TP/(TP+FN) | Of actual positives, how many were caught | When misses are costly |
| F1 | 2·P·R/(P+R) | Harmonic mean of P and R | When both matter |
| Fβ | (1+β²)·P·R/(β²·P+R) | Give R β× weight | When favoring one side |
| AUC / AP | See below | Ranking ability / average precision for positives | Threshold-independent overall comparison |
python
print(classification_report(y_test, y_pred))
print(f"F1 = {f1_score(y_test, y_pred):.3f}")classification_report gives P/R/F1 and per-class sample counts in one go; it's the main tool for daily iteration.
3. Threshold Is Part of Evaluation, Not a Model Parameter
Most classifiers output probabilities, not labels; predict() just cuts at 0.5. When business costs are asymmetric (e.g., missing a diagnosis is far worse than a false alarm), 0.5 is often not the optimal cut point. The right posture: treat the threshold as a hyperparameter to scan:
python
for threshold in [0.3, 0.4, 0.5, 0.6, 0.7]:
y_t = (y_prob >= threshold).astype(int)
print(f"Threshold {threshold}: Precision={precision_score(y_test, y_t):.3f} "
f"Recall={recall_score(y_test, y_t):.3f} "
f"F1={f1_score(y_test, y_t):.3f}")Which threshold to choose depends on the business cost table: one missed diagnosis = how many yuan, one false alarm = how many yuan; find the cut point that minimizes total cost. This is the key step of "moving evaluation from inside the model to the business layer" — evaluation output should be a "threshold-metrics-cost" table, not a label array.
4. PR and ROC Curves: Plotting the Model's Full Spectrum
ROC looks at the trade-off between true positive rate and false positive rate at all thresholds; the PR curve looks at the trade-off between precision and recall at all thresholds. AUC is the area under the curve, measuring "ranking ability": the probability that, given a random positive and a random negative, the model gives the positive a higher score.
python
# ROC curve and AUC
fpr, tpr, roc_thresholds = roc_curve(y_test, y_prob)
roc_auc = roc_auc_score(y_test, y_prob)
# PR curve and AP (Average Precision, area under PR curve)
precision, recall, pr_thresholds = precision_recall_curve(y_test, y_prob)
ap = average_precision_score(y_test, y_prob)
print(f"ROC-AUC = {roc_auc:.3f} PR-AUC(AP) = {ap:.3f}")
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 4))
ax1.plot(fpr, tpr, label=f"AUC={roc_auc:.3f}")
ax1.plot([0, 1], [0, 1], "k--", label="Random baseline")
ax1.set(xlabel="FPR", ylabel="TPR", title="ROC Curve")
ax1.legend()
ax2.plot(recall, precision, label=f"AP={ap:.3f}")
ax2.axhline(y_train.mean(), color="k", ls="--",
label="Positive ratio baseline")
ax2.set(xlabel="Recall", ylabel="Precision",
title="PR Curve")
ax2.legend()
plt.tight_layout()
plt.show()Note that ROC's random baseline is the y=x diagonal line (independent of positive ratio), while PR curve's baseline is the horizontal line at the positive ratio. When positives are only 10%, PR curve's baseline is 0.1 — this is exactly the key to "ROC lies" in Section VI.
5. Package Everything Above Into One Function
Daily iteration needs "one command output full evaluation." Consolidating scattered code into a reusable evaluator is the first step of evaluation engineering:
python
def evaluate_binary_classifier(y_true, y_prob, name="Model"):
"""Output a full offline evaluation report for a binary classifier."""
results = {}
results["ROC-AUC"] = roc_auc_score(y_true, y_prob)
results["AP"] = average_precision_score(y_true, y_prob)
# Scan thresholds, find P/R for best F1
p, r, th = precision_recall_curve(y_true, y_prob)
f1s = 2 * p * r / (p + r + 1e-9)
best = np.argmax(f1s)
results["best_threshold"] = th[best]
results["best_P"] = p[best]
results["best_R"] = r[best]
results["best_F1"] = f1s[best]
results["confusion_matrix"] = confusion_matrix(
y_true, (y_prob >= results["best_threshold"]).astype(int))
print(f"===== {name} =====")
for k, v in results.items():
if k != "confusion_matrix":
print(f"{k}: {v:.3f}")
print("Confusion matrix (at best F1 threshold):\n", results["confusion_matrix"])
return resultsA clear evaluation function should be output information volume > input code volume, and always hangable into regression tests.
III. Regression Evaluation: More Than Just MSE
Regression evaluation seems simple on the surface (one error number), but a single error number is exactly the biggest pitfall — MSE is hijacked by outliers, R² is deceived by variance. The approach: metrics as foundation + residual diagnosis.
1. Three Core Metrics + One Yes/No Question
| Metric | Formula | Property | When It Dominates |
|---|---|---|---|
| MSE | (1/n)Σ(yᵢ-ŷᵢ)² | Penalizes large errors (square amplifies) | Large errors unacceptable |
| RMSE | √MSE | Same units as y, intuitive | Need to compare directly with y |
| MAE | (1/n)Σ|yᵢ-ŷᵢ| | Robust to outliers | Error distribution has heavy tails |
| R² | 1 − SS_res/SS_tot | Explained variance relative to mean baseline | Relative comparison between models |
A key yes/no question: when outliers come from "normal but extreme business" (large orders, extreme weather), use MAE; when outliers represent "system failures / dirty data," clean first rather than switching metrics.
python
from sklearn.datasets import make_regression
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score
X, y = make_regression(n_samples=3000, n_features=10, noise=15, random_state=42)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=42)
reg = RandomForestRegressor(n_estimators=200, random_state=42)
reg.fit(X_tr, y_tr)
y_hat = reg.predict(X_te)
mse = mean_squared_error(y_te, y_hat)
rmse = np.sqrt(mse)
mae = mean_absolute_error(y_te, y_hat)
r2 = r2_score(y_te, y_hat)
print(f"MSE={mse:.1f} RMSE={rmse:.1f} MAE={mae:.1f} R²={r2:.3f}")If RMSE is clearly larger than MAE (e.g., 2×), the error distribution's right tail is heavy — a few samples contribute most of the error, and this is exactly where the next step, residual analysis, digs.
2. Residual Analysis: Evaluation's X-Ray
Residuals = true value − predicted value. A good model's residuals should have mean 0, stable variance, and no relationship with features or predictions. Four plots are enough:
python
residuals = y_te - y_hat
fig, axes = plt.subplots(2, 2, figsize=(11, 8))
# (1) Residuals vs predicted: should be a random band, no trumpet / bend shape
axes[0, 0].scatter(y_hat, residuals, s=8, alpha=0.5)
axes[0, 0].axhline(0, color="r", lw=1)
axes[0, 0].set(xlabel="Predicted", ylabel="Residual", title="Residuals vs Predicted")
# (2) Residual histogram: should be approximately normal, centered at 0
axes[0, 1].hist(residuals, bins=50, alpha=0.7)
axes[0, 1].axvline(0, color="r", lw=1)
axes[0, 1].set(xlabel="Residual", ylabel="Count", title="Residual Distribution")
# (3) Residuals vs most important feature: check for missed non-linear relationships
imp = np.argsort(-reg.feature_importances_)[0]
axes[1, 0].scatter(X_te[:, imp], residuals, s=8, alpha=0.5)
axes[1, 0].axhline(0, color="r", lw=1)
axes[1, 0].set(xlabel=f"Most important feature #{imp}", ylabel="Residual",
title="Residuals vs Most Important Feature")
# (4) Cumulative residual ratio: see what proportion of samples contribute most error
axes[1, 1].plot(np.sort(np.abs(residuals))[::-1].cumsum()
/ np.abs(residuals).sum())
axes[1, 1].set(xlabel="Sample index (sorted by error descending)", ylabel="Cumulative error ratio",
title="Error Concentration Curve")
plt.tight_layout()
plt.show()Four typical symptoms and corresponding prescriptions:
| Residual Symptom | Meaning | Action |
|---|---|---|
| Trumpet shape (residuals increase with predicted value) | Heteroscedasticity, possibly missing log transform | Model y after log |
| Bent / opening curve | Missed non-linear terms or feature interactions | Add feature interactions, switch to tree models |
| Residual mean significantly deviates from 0 | Systematic bias (underfitting or features leak directionally) | Check features, increase capacity |
| Error concentrates in few samples | Heavy tails, a few "large orders" dominate error | Fall back on MAE for evaluation, chase these 5% separately |
High R² but systematic patterns in residuals means the model is "accurate on the surface, biased inside." The essence of regression evaluation isn't looking at one score; it's making the misalignment between model and data visible. How bias-variance and regularization further affect errors, see Overfitting and Regularization.
IV. The Right Way to Use Cross-Validation
A single train/test split result has randomness — change the seed, and AUC might jump from 0.83 to 0.87. Cross-validation (CV) uses the mean ± std of multiple splits to give a stable estimate, while using every sample.
1. StratifiedKFold: The Standard Choice for Classification
python
from sklearn.model_selection import StratifiedKFold, cross_val_score
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(model, X, y, cv=skf, scoring="roc_auc")
print(f"5-fold ROC-AUC: {scores.mean():.3f} ± {scores.std():.3f}")Two key points:
- shuffle=True + random_state=42: without shuffling, if data is ordered by class, each fold's class distribution will be severely imbalanced;
- Stratified rather than plain KFold: ensures each fold's positive/negative ratio matches the full data, especially critical on small, imbalanced datasets.
cross_val_score is for quick estimation, but to get per-fold prediction probabilities for error analysis, loop manually:
python
oof = np.zeros(len(X)) # out-of-fold prediction probabilities
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for fold, (tr_idx, va_idx) in enumerate(skf.split(X, y)):
clf = RandomForestClassifier(n_estimators=200, random_state=42)
clf.fit(X[tr_idx], y[tr_idx])
oof[va_idx] = clf.predict_proba(X[va_idx])[:, 1]
print(f"OOF AUC = {roc_auc_score(y, oof):.3f}")"Out-of-fold predictions" (OOF) is the golden data source for Section V error analysis: each sample's prediction comes from a model that never saw it, so it's safe to use for error pattern statistics.
2. Leakage: CV's Most Concealed Killer
CV's statistical validity rests on one premise: each fold's training data contains only information from that fold's training set. The three most common leakages:
| Leakage Source | Example | Consequence |
|---|---|---|
| Preprocessing done outside CV | Standardize / mean-impute on full data first, then do CV | Validation set information leaks into training, inflated metrics |
| Feature selection leaks | Select features on full data before entering CV | Selected features that "leak the future" |
| Random split for time data | Use day 3 data to train, day 1 data to validate | Model "sees the future," time-series scenario fails |
The standard solution is Pipeline — let each preprocessing step be inside CV, re-fit per fold:
python
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([
("scaler", StandardScaler()), # re-fit inside each fold
("clf", LogisticRegression(max_iter=2000)),
])
scores = cross_val_score(pipe, X, y, cv=skf, scoring="roc_auc")
print(f"Pipeline 5-fold AUC: {scores.mean():.3f} ± {scores.std():.3f}")3. Group and Time-Series: When Samples Aren't Independent
If your data has group dependency structure (multiple behaviors from the same user, multiple orders from the same batch of goods, multiple patches from the same image), plain KFold splits the same group's data across train and validation — the model has "seen the same user," so evaluation is naturally inflated. Use GroupKFold to guarantee the same group's samples always stay in the same fold:
python
from sklearn.model_selection import GroupKFold
# group array: which user/batch each sample belongs to
group = np.arange(len(X)) // 10 # example: every 10 samples from the same user
gkf = GroupKFold(n_splits=5)
scores = cross_val_score(model, X, y, cv=gkf, groups=group, scoring="roc_auc")
print(f"GroupKFold AUC: {scores.mean():.3f} ± {scores.std():.3f}")Time-series data uses TimeSeriesSplit (only "past to train, future to validate"), and note two rules: training always predates validation; validation must not contain future information from the training set (e.g., when features include "historical mean," compute them point-by-point in a rolling manner, not all at once).
4. Nested CV: When CV Is Used for Both Tuning and Evaluation
If you use CV for both tuning (GridSearchCV) and reporting final metrics, the second set of scores is contaminated by the tuning process (the validation set has been seen too many times). The rigorous approach is nested CV: outer CV estimates generalization, inner CV handles tuning.
python
from sklearn.model_selection import GridSearchCV
param_grid = {"n_estimators": [50, 200], "max_depth": [5, None]}
inner = GridSearchCV(RandomForestClassifier(random_state=42),
param_grid, cv=StratifiedKFold(3, shuffle=True, random_state=1))
outer_scores = cross_val_score(inner, X, y,
cv=StratifiedKFold(5, shuffle=True, random_state=2),
scoring="roc_auc")
print(f"Nested CV AUC: {outer_scores.mean():.3f} ± {outer_scores.std():.3f}")Nested CV doubles computation; you don't need it every daily iteration, but at key decision points ("what is this model truly worth"), running it once is worth it. The trade-off between CV and bootstrap for small samples, refer to Kohavi's classic study (see references).
V. Error Analysis and the Iteration Loop
Evaluation isn't the end; it's the input for the next iteration. Without error analysis, evaluation is just "reporting a score"; with it, evaluation becomes "finding where the model improves next."
1. What Error Analysis Is
In one sentence: on the validation set (or OOF predictions), find the samples the model got wrong, understand why, and based on this decide the next action. The core tool is "turn evaluation results into a queryable table."
python
# Turn the test set into an analyzable DataFrame
df_results = pd.DataFrame({
"y_true": y_test,
"y_prob": y_prob,
"y_pred": y_pred,
"is_error": y_test != y_pred,
"abs_error": np.abs(y_test - y_pred), # for regression
"confidence": np.where(
y_prob >= 0.5, y_prob, 1 - y_prob), # model confidence
})2. Group Statistics by Error Type
With the table, error analysis is a series of groupby operations. Below is a directly usable grouping matrix:
python
# ① By business field: which channel/time period/customer segment has the highest error rate?
df_results["group"] = X_test[:, 0] > 0 # example: split by a feature into two groups
group_stats = (df_results.groupby("group")["is_error"]
.agg(errors="sum", count="count", error_rate="mean"))
print(group_stats)
# ② By confidence tier: what proportion of errors are "model was very confident and wrong"?
df_results["conf_bin"] = pd.cut(df_results["confidence"],
bins=[0, 0.6, 0.8, 0.9, 1.0])
print(df_results.groupby("conf_bin", observed=True)["is_error"]
.agg(error_rate="mean", count="count"))
# ③ Average confidence of error samples vs correct samples
print(df_results.groupby("is_error")["confidence"].mean())Three typical findings and their directed actions:
| Analysis Finding | Implied Root Cause | Action |
|---|---|---|
| One segment / channel has error rate far above average | This subset lacks features or samples | Add features, add data for this subset |
| Error samples all have high confidence | Model overconfident / poorly calibrated | Use CalibratedClassifierCV for probability calibration |
| Error samples happen to be "rare sub-classes" | Class imbalance / data sparsity | See Section VI's resampling and weighting approaches |
| Error samples share a common time period | Time-correlated drift | Add time features, switch to time-based split |
The right way to do error analysis
Don't just look at summary numbers; personally flip through 20–30 raw error samples yourself. Numbers tell you "where errors are concentrated," samples tell you "why the errors are absurd" — e.g., you thought the model learned semantics, but it's actually copying repeated sentences from the training set. This cognitive gap only appears when you see raw samples. Interpretability tools (SHAP, partial dependence plots) can turn this "manual sample-flipping" into systematic attribution; see Interpretability and Fairness.
3. Turning Analysis Into Action: The Evaluation–Iteration Loop
Error analysis produces hypotheses, not conclusions. Each analysis converges on a "finding → hypothesis → action" table; only do the one with the biggest impact, then re-evaluate:
Evaluate (fixed protocol) → Error analysis (group stats + flip samples)
↑ ↓
Re-evaluate, compare to previous round Form hypothesis
↑ ↓
Record: changed data / features / model / threshold → Execute one minimal changeA practical iteration discipline:
- Change one variable at a time: add features, switch models, adjust thresholds — do them separately, or improvement attribution is unclear;
- Every iteration has a comparison record: what changed, before/after val metrics, whether new errors were introduced;
- Calibrate changes back to "business goals": val AUC went up 0.01, but the online conversion proxy metric didn't move → the improvement is ineffective;
- Acknowledge "evaluation saturation": if two to three rounds produce no significant gains, stop and do a big review, rather than continuing to fine-tune.
The full skeleton of this loop can be mapped onto the workflow in Building an ML Project from Scratch; the most common mistakes in iteration (repeatedly tuning on validation set, metrics decoupled from business, etc.) have a ready-made avoidance list in Common Pitfalls and Anti-Patterns.
VI. Class Imbalance Evaluation: Why PR-AUC Usually Beats ROC
Imbalance isn't "data has a problem"; it's the default state of most real businesses: fraud, disease, clicks, failures — rare events always occupy the minority. Under imbalance, the choice of evaluation metric directly determines whether you see the truth or an illusion.
1. Why ROC Can "Lie"
ROC's horizontal and vertical axes are TPR and FPR, both of which are ratios with their respective classes as denominators — it treats positive and negative classes symmetrically, so ROC shape changes very little whether positives are 1% or 50%. But PR curve's vertical axis is precision, whose baseline is the positive ratio: when positives are 1%, random guessing gives 0.5 ROC-AUC, but PR baseline is only 0.01.
Consider an extreme but real example (positives at 1%, 10,000 test samples):
| Scenario | ROC-AUC | AP (PR-AUC) |
|---|---|---|
| Model predicts all samples negative | 0.5 (equal to random) | Undefined (no positive predictions) |
| Model catches 60% positives, 50 false alarms | Can reach 0.98+ | ~0.55 |
| Model catches 10% positives, 3 false alarms | Can also reach 0.95 | ~0.25 |
In the same experiment, ROC-AUC looks pleasantly high, but AP is shockingly low — because ROC is "diluted" by the huge negative class, while the PR curve directly reflects "when you say it's positive, how likely are you right?" The lower the positive ratio, the more ROC and PR conclusions diverge. For rigor, recommend looking at both, but when the business cares about "are positives found correctly," take PR-AUC/AP as the standard. Theoretical details are in Davis & Goadrich and Saito & Rehmsmeier's papers (references).
2. The Correct Posture for Imbalanced Evaluation
- Report PR-AUC/AP, simultaneously print confusion matrix absolute values (absolute TP/FN/FP out of 10k are more business-meaningful than ratios);
- Use Fβ instead of F1: increase β when misses are costly;
- Report P/R pairs at business thresholds: don't chase "best F1"; instead select threshold from a cost table, then report P/R;
- Don't just look at accuracy — it's dominated by the 99% majority class.
3. Mitigation Approaches, In Order
Fix evaluation first, then data, then model only last. Get the order wrong, and everything upstream is distorted:
python
# ① Don't change data, first weight the loss function (internal model solution)
from sklearn.ensemble import RandomForestClassifier
weighted = RandomForestClassifier(n_estimators=200,
class_weight="balanced", random_state=42)
weighted.fit(X_train, y_train)
print("class_weight=balanced AP:",
average_precision_score(y_test, weighted.predict_proba(X_test)[:, 1]))
# ② Oversample minority class (resampling approach, must be done inside CV)
from imblearn.over_sampling import SMOTE
from imblearn.pipeline import Pipeline as ImbPipeline
smote_pipe = ImbPipeline([
("smote", SMOTE(random_state=42)),
("clf", RandomForestClassifier(n_estimators=200, random_state=42)),
])
smote_scores = cross_val_score(smote_pipe, X, y, cv=skf, scoring="average_precision")
print(f"SMOTE 5-fold AP: {smote_scores.mean():.3f} ± {smote_scores.std():.3f}")Resampling must never appear in the validation pipeline
SMOTE or undersampling can only happen inside training folds — if you SMOTE on full data first, then split, synthesized minority samples appear in both train and validation, and validation metrics are inflated by their own "duplicates." The code above uses imblearn.pipeline.Pipeline specifically to guarantee SMOTE executes per-fold inside CV.
Beyond resampling, there's collecting more positive samples (always most effective), changing the problem to anomaly detection / ranking (when positives are single digits), and cost-sensitive learning (writing false alarm/miss costs directly into the objective function). Remember: data itself being imbalanced isn't the disease; the evaluation approach being wrong is the disease — class_weight and threshold selection are often more effective than fancy resampling.
VII. LLM Application Evaluation: A Brief Intro
LLM applications (RAG, Agents, summarization, code generation) evaluation differs fundamentally from classical ML: outputs are open-ended text, correctness is highly subjective, and errors aren't "right/wrong" but "which part is wrong." But the four steps of evaluation system design are fully applicable — just the implementation differs. This section is a quick survey; full solutions are in Large Language Model Application Evaluation.
1. Constructing the Eval Set
LLM evaluation also needs "train / dev / test" layering, but what's split here isn't model weights, but prompts and system design:
| Set | Purpose | Recommended Size |
|---|---|---|
| Dev set | Iterate prompts, tune RAG parameters | 50–200 samples |
| Test set | Final acceptance, ideally close to real distribution | 200–1000 samples |
| Golden set | Auto-regression, fixed unchanged | 100–500 samples |
Collection priority: online real requests > manually constructed typical scenarios > mined failure samples from user feedback. A common mistake is the eval set only containing "smooth questions," so the model can't even handle "a user asked an irrelevant question" — the eval set must deliberately include edge cases (ambiguity, missing context, long-tail dialects, malicious input).
2. Human Evaluation: The Only Trustworthy Scale
Automated metrics still can't replace humans in judging "is this summary accurate, is this response offensive?" To land a reproducible human evaluation:
- Define a scorecard: correctness, faithfulness (faithfulness to context), usefulness, style/tone, each 1–5 points with clear anchors;
- Multiple people score, compute inter-annotator agreement; re-discuss rules for samples with disagreement;
- Re-score the same batch after each change, avoiding "different batches, different standards" drift.
Human evaluation is slow and expensive, so its position is as the calibration ruler for automated metrics, not the main tool for daily iteration.
3. Automated Metrics: Can Save Manpower, But Don't Worship
| Type | Examples | Purpose | Limitation |
|---|---|---|---|
| Text similarity | BLEU, ROUGE | Summarization, translation | Completely ineffective for "semantically correct but differently worded" |
| Component-level metrics | RAG retrieval Recall@k, hit rate | Evaluate RAG retrieval layer in isolation | Requires layered evaluation design |
| LLM-as-judge | Use a strong model (e.g., GPT-4) to score a weaker model | Faithfulness, relevance | Self-preference bias |
Three landing key points for LLM-as-judge: provide structured scoring criteria (far more stable than free-form comments); run judge consistency checks (score a batch of samples with both humans and judges; if correlation is below threshold, it's unreliable); prioritize scoring dimensions with anchors like "factuality / relevance" over "writing quality." More nuanced issues on hallucination and eval bias are in the Large Language Models deep-dive.
4. Layered Evaluation: RAG Applications Should Be Split Into at Least Three Layers
For an end-to-end RAG Q&A system, when the overall metric is poor, you have no idea whether retrieval didn't retrieve or the model didn't answer correctly. Evaluate each layer separately:
User question ──► Retrieval layer (Recall@k: was the correct document recalled?)
│
▼
Generation layer (Faithfulness: does the answer faithful to the retrieved docs?)
│
▼
Interaction layer (Usefulness: did the answer solve the user's problem?)Each layer uses different metrics, different failure modes, and different optimization actions. The value of a single end-to-end score is far less than this three-layer health check table.
VIII. Evaluation Engineering: Making Evaluation a Systematic Part
Doing evaluation once isn't hard; the challenge is automating and reproducing evaluation after every change, and continuously monitoring post-deployment. This belongs to the realm of MLOps and Production Deployment; here are three minimum must-do actions.
1. Evaluation Scripts In Repo, Same Repo as Model Code
Evaluation isn't a casual brush in Jupyter; it's versioned code. Recommended directory structure:
project/
├── model/ # training code
├── evals/
│ ├── metrics.py # metrics and evaluation functions (e.g., evaluate_binary_classifier from Section II)
│ ├── datasets/ # test sets, golden sets (with version numbers)
│ ├── run_eval.py # one-click eval entry, outputs JSON/Markdown reports
│ └── tests/ # unit tests for the evaluation itself
└── reports/ # historical evaluation reports, traceablerun_eval.py outputs a report with metrics, data version, code commit, random seed, run time every time — this is the only evidence chain for "what changed last round, and did it improve?"
2. Use Golden Sets for Regression Testing
Turn evaluation from "run when someone remembers" to "run automatically after every change":
- Pick 100–500 golden samples; after evaluation, assert key metrics don't fall below threshold (e.g.,
assert ap >= 0.40); - Integrate into CI: changing model code, features, or data scripts all trigger evaluation; metrics below threshold → build fails;
- Make "never regress" an engineering convention: models before deployment must pass all golden assertions, and not fall below the current online model's metrics on the test set.
python
# Skeleton of run_eval.py (excerpt)
def check_regression(results, thresholds):
failures = []
for metric, limit in thresholds.items():
if results[metric] < limit:
failures.append(f"{metric}={results[metric]:.3f} < {limit}")
if failures:
raise RuntimeError("Regression test failed: " + "; ".join(failures))3. Post-Deployment Monitoring: From "One-Time" to "Continuous"
The premise of offline evaluation is "test set represents the future," but distributions drift online. Deploy at least three monitoring channels:
| Monitoring Item | Method | Warning Signal |
|---|---|---|
| Data drift | Feature distribution (PSI / KL divergence) daily comparison | PSI > 0.2 needs review |
| Label drift | Online label / result distribution changes | Sudden shift in positive rate |
| Business proxy metrics | Online conversion rate, success rate, user complaints | Metrics diverge from offline expectations |
Complete practices for drift detection and model retraining cycles are in MLOps and Production Deployment. One principle running throughout: online monitoring metrics must be the same as to offline evaluation metrics — otherwise you'll simultaneously have two mutually contradictory truths.
IX. Trade-offs and Considerations
There's no free lunch in evaluation systems; every bit of rigor has a cost. These trade-offs should be made proactively and documented, not passively borne.
- Metric fidelity vs. iteration speed: nested CV, repeated experiments, thousand-person scoring are the most rigorous, but each iteration takes hours or days. Use lightweight evaluation (single StratifiedKFold + proxy metrics) for daily iteration; use heavy evaluation (nested CV, human scoring) at key decision points (deployment, architecture changes). Two-tier evaluation is a common solution balancing both.
- Test set size vs. estimation precision: larger test sets = more stable metrics, but less data for training. When data is limited, prefer to train set slightly smaller to keep the test set representative — the test set is the "final judge," and judges can't starve.
- Metric granularity vs. interpretability: NDCG, KS, calibration error are more precise, but business stakeholders don't understand them; accuracy is universally understood but often misleading. Put both "coarse metrics the business understands" and "precise metrics engineers use" in the report, don't pick one.
- Automation vs. human review: automated metrics are cheap but may be distorted (especially for LLM evaluation); human evaluation is credible but expensive. The compromise is automated screening + human spot-checking: automation runs the full set, sample by layer ("predicted failure / low confidence / edge cases") for human review.
- The temptation to use validation as test set: after 50 iterations, the "validation set" has actually been seen by you 50 times. A honest evaluation system reserves an untouched final test set, and does one-time final confirmation after validation metrics saturate. See Common Traps in Evaluation.
A final overarching principle: the value of an evaluation system doesn't depend on how advanced its metrics are, but on whether it can stably answer "will my business improve after this model goes live?" Metrics are proxies; business is reality; evaluation's purpose is to make the gap between the two as small as possible, and known.
X. Further Reading
- Continue on-site: theory behind metrics is in Model Evaluation and Validation; evaluation connected to tuning is in Hyperparameter Tuning Practice; evaluation connected to production systems is in MLOps and Production Deployment; complete project workflow is in Building an ML Project from Scratch; error point summaries are in Common Pitfalls and Anti-Patterns; LLM deep-dive is in Large Language Models.
- Fill basics first: this article assumes you've understood What is Machine Learning; look up terminology anytime at Glossary; the data layer (collection, cleaning, leakage sources) is in Data Engineering Fundamentals.
References
- scikit-learn: Metrics and scoring: quantifying the quality of predictions — definitions, formulas, and call methods for all metrics, the official basis for this article's code
- scikit-learn: Cross-validation: evaluating estimator performance — KFold / StratifiedKFold / GroupKFold / TimeSeriesSplit / nested CV official docs
- Hastie, Tibshirani, Friedman. The Elements of Statistical Learning — authoritative textbook on model evaluation, bias-variance, and error estimation (Chapter 7)
- James, Witten, Hastie, Tibshirani. An Introduction to Statistical Learning (ISLR) — introductory authoritative textbook on classification/regression evaluation and cross-validation (Chapters 2, 5)
- Fawcett. An Introduction to ROC Analysis (Pattern Recognition Letters, 2006) — classic review of ROC curves
- Davis & Goadrich. The Relationship Between Precision-Recall and ROC Curves (ICML 2006) — formal proof of PR vs ROC relationships, theoretical basis for imbalanced evaluation
- Saito & Rehmsmeier. The Precision-Recall Plot Is More Informative than the ROC Plot when Evaluating Binary Classifiers on Imbalanced Datasets (PLoS ONE, 2015) — empirical study showing PR beats ROC for imbalanced datasets
- Kohavi. A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection (IJCAI 1995) — classic comparison of cross-validation and bootstrap estimation
- imbalanced-learn official docs — SMOTE and other resampling methods and imblearn Pipeline official basis
- Zheng et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (NeurIPS 2023) — bias analysis and validation methods for LLM-as-judge
- Papineni, Roukos, Ward, Zhu. BLEU: a Method for Automatic Evaluation of Machine Translation (ACL 2002) — original BLEU paper
- Lin. ROUGE: A Package for Automatic Evaluation of Summaries (ACL 2004) — original ROUGE paper