Theme
Deep Learning Evaluation and Experiments
One-line definition: Evaluation is the only way to answer "how good is the model?" and experimental discipline determines whether the evaluation is trustworthy. Training loss tells you "has the model memorized the data?"; evaluation metrics tell you "can the model generalize?" The gap between the two — and how to bridge it — is where Loss Functions and Output Layers and this article each play their part.
I. Train/Validation/Test Splits and the Data Leakage Defense Line
The standard approach is to split the data into three sets:
- Train set: Visible during training, used to update parameters.
- Validation set: Repeatedly used during training — for hyperparameter selection, early stopping, model selection. Because it has been used dozens or hundreds of times, it is already "contaminated" and no longer represents true generalization.
- Test set: Touch it only once. After the final model is locked in, use it to report a score, and then never go back to tune based on it.
Data leakage — when "test information seeps into the training process" — is the #1 cause of evaluation failure. Common leakage patterns:
- Preprocessing done on the full dataset: Normalizing using global statistics (mean/variance computed on the test set), or fitting standardization, PCA, and imputation on the full data. These must be fitted on the training set, then applied to the validation and test sets.
- Homogeneous samples across sets: The same user's data or frames from the same video end up in both train and test — the evaluation looks artificially inflated. Split by user/session/time (see the Data Engineering section).
- Peeking at the test set during hyperparameter tuning: Using the test set repeatedly turns it into a validation set.
Symptoms of Leakage
Test scores that are abnormally high, or "validation going up and test skyrocketing" to an unreasonable degree — suspect leakage first, model brilliance second. See the full defense-line checklist in the "Types of Data Leakage and Defense Lines" section of the Data Engineering chapter.
II. Overfitting/Underfitting Diagnostics: Learning Curves
Learning curves plot "training loss / validation loss over training epochs (or data volume)" and let you diagnose the state at a glance:
- Underfitting: Both training and validation losses are high and decline slowly — the model lacks capacity or hasn't been trained enough. Countermeasures: increase capacity, increase training duration. See the capacity discussion in Neural Networks Fundamentals.
- Overfitting: Training loss keeps dropping while validation loss first drops then rises — the model is memorizing the training set. Countermeasures: regularization, early stopping, data augmentation, reduce capacity. See Overfitting and Regularization.
- Ideal: Both curves are low and converge with a small gap — good generalization.
Early stopping (at the point where validation loss starts rising) is the highest ROI regularization technique. Note: the validation set itself may have been evaluated repeatedly for early stopping — so the final score is always settled on the test set.
III. Classification Metrics: Accuracy, Precision, Recall, F1, AUC, AP
First, define the four cells of the confusion matrix: TP (true positive), FP (false positive), FN (false negative), TN (true negative).
| Metric | Formula | Question It Answers |
|---|---|---|
| Accuracy | (TP+TN)/(all) | Proportion of overall correct predictions |
| Precision | TP/(TP+FP) | Of what was predicted "positive," how many were actually positive? |
| Recall | TP/(TP+FN) | Of all actual positives, how many were retrieved? (how many missed?) |
| F1 | 2·P·R/(P+R) | Harmonic mean of P and R, balancing both |
| AUC | Area under the ROC curve | Probability that, given a random positive and a random negative, the model scores the positive higher |
| AP | Area under the PR curve | Integration of P over R at different thresholds; commonly used in detection tasks |
Accuracy is misleading under class imbalance: on a dataset that is 99% negative, always predicting "negative" gives 99% accuracy — the model is worthless. In such cases, look at precision/recall/F1, AUC, or switch to weighted/resampled training (paired with Focal Loss from Loss Functions and Output Layers).
Three rules of thumb for metric selection:
- Don't use accuracy when costs are asymmetric: If a false negative (e.g., medical screening miss) is costly, look at recall; if a false positive (e.g., spam misclassification) is costly, look at precision.
- AUC only measures ranking ability, not the calibration of predicted probabilities. When you need "accurate probabilities," look at Brier score or calibration curves.
- Detection/retrieval tasks (multiple bounding boxes, multiple documents) are better evaluated with AP/mAP. For more deployment-specific details, see Evaluation in Practice.
IV. Regression Metrics: MSE, MAE, R²
- MSE: Penalizes large errors, sensitive to outliers, units are squared.
- MAE: Robust, units match the original scale, more interpretable.
- R² (coefficient of determination):
1 − SSE/SST, explains "how much better the model is compared to a mean baseline." R² = 0 is equivalent to "all predictions are just the mean"; negative values mean the model is worse than the mean baseline. It is dimensionless and comparable across tasks.
The principle of regression evaluation: the metrics must align with the loss and business objectives. If you train with Huber (see the Loss Functions section) but only report MSE for evaluation, you will mask outlier problems — first figure out "which type of error hurts most in the business context."
V. Generative Model Evaluation: FID, IS, NLL
Evaluating generative models (GANs, diffusion models, LLMs) is harder: there is no "right or wrong" ground truth. Mainstream metrics:
- NLL (Negative Log-Likelihood): The average probability the model assigns to real data. For explicit density models (VAEs, autoregressive models), it can be computed directly and is one of the gold standards for "fit quality." See Generative Models.
- IS (Inception Score, 2016): Uses an Inception-v3 classifier to measure whether "generated samples are clear (low class entropy) and diverse (high class distribution entropy)." Weakness: only measures in-distribution classes and is not sensitive to mode collapse.
- FID (Fréchet Inception Distance, 2017): Compares the distance between two Gaussian distributions of real and generated images in the Inception feature space (mean + covariance). More sensitive to mode collapse and blur, making it the de facto standard for image generation. Note: FID has its own bias — it is unstable when the sample count is too low.
- Human evaluation: LLM evaluation often falls to human scoring and preference (RLHF uses human preferences; see Deep Reinforcement Learning). It is expensive and hard to reproduce, but remains irreplaceable to date.
Principles of Generative Evaluation
Cross-validate with multiple metrics: a single metric can always be "gamed." For image generation, report FID + IS + manual inspection; for text generation, report perplexity + human evaluation + downstream task performance. For metric selection and specific details, see Evaluation in Practice.
VI. Hyperparameter Search and Experiment Management
Three mainstream approaches for searching hyperparameters (learning rate, batch size, number of layers, dropout probability, etc.):
- Grid search: Cartesian product enumeration. Simple, but explodes when dimensions go beyond 3. Not recommended for more than 3 dimensions.
- Random search: Randomly sample from prior distributions for each hyperparameter. Bergstra & Bengio proved in 2012 that random search is more efficient than grid search under the same budget — because many hyperparameters have their impact dominated by "a few key parameters," and random sampling explores key dimensions more thoroughly.
- Bayesian optimization: Uses surrogate models like Gaussian processes to guide the search, densifying sampling in key regions. Representative tools: Optuna, Weights & Biases Sweeps. Most cost-effective when dimensions are high and individual training runs are expensive.
The fundamentals of experiment management (without these, credible research is out of reach):
- Fix the random seed: In PyTorch, set
torch.manual_seed+random.seed+numpy.random.seed, and disable nondeterministic operators (or record their impact). - Record everything: Code version, data version, hyperparameters, environment, seed, and per-epoch metrics. Infrastructure for automation is covered in MLOps and Model Deployment.
- Controlled comparisons: Change only one variable, freeze everything else. This is the foundation of all credible conclusions.
VII. Baseline Discipline
Run a simple, working baseline first, then talk about innovation. Three hard rules:
- Always have a naive baseline: for classification, the majority class; for regression, mean prediction; for time series, the previous-value predictor. If your model can't beat "always predict the majority class," there is something wrong with the evaluation or the data itself.
- Run published models as a reference: reproduce a public model from a standard library (e.g., ResNet, BERT open-source weights) and align with the paper's numbers on the same data and metric — the alignment process itself checks whether your implementation is correct.
- Start small, then scale up: first get the pipeline working on a small dataset with a small model (see Incremental Tutorial), then scale up. Once you scale up, debugging specific changes becomes very difficult.
VIII. Common Evaluation Pitfalls
- Only reporting the best single run, without mean ± variance: train multiple times with different seeds and report mean and variance. A single result could be luck.
- Repeatedly using the test set: if you tuned hyperparameters by peeking at the test set, the "test set" no longer exists.
- Cherry-picking metrics: F1 isn't good? Report accuracy. Accuracy isn't good? Report AUC — pre-register the primary metric.
- Ignoring computational overhead: report FLOPS/latency/parameter count. Without these, "better" has nothing to compare against. Deployment metrics are covered in MLOps and Model Deployment.
- Ignoring out-of-distribution (OOD) scenarios: the test set shares the same distribution as the training set, but real deployment encounters new distributions. See the Data Engineering section for data splitting and distribution drift discussions.
Tradeoffs
Validation set size vs. training data: larger validation sets make hyperparameter selection more reliable but leave less for training. Rule of thumb: for image classification with < 1000 classes, 5%–10% for validation is fine; for data-scarce scenarios, use cross-validation (at the cost of k× training runs).
Metric comprehensiveness vs. interpretability: comprehensive metrics like AUC/FID are easy to compare but hard to diagnose; per-class metrics and bad-case sampling can pinpoint issues but are fragmented information. Be comprehensive first, then diagnostic — one primary metric, many diagnostic metrics.
Experimental rigor vs. iteration speed: recording everything slows you down, but not recording means you didn't do it. A compromise: solidify a "reproducible template" into scripts, so running experiments is just filling in parameters. See Training Recipes and Hyperparameter Tuning.
Evaluation determines project direction: do we need more data? Do we need a different model? Is it ready to deploy? Getting evaluation right is more valuable than piling on more models. Common evaluation libraries and public benchmarks are covered in Datasets and Tools Archive; evaluation-related terminology is in Glossary.
Further Reading
- Optimization and Gradient Descent — the optimization perspective on hyperparameters and training recipes
- Overfitting and Regularization — how to treat problems identified by learning curves
- Common Pitfalls and Anti-patterns — common errors in evaluation and experiments
- Datasets and Tools Archive — public benchmarks and evaluation tools
- Build a Deep Learning Project from Scratch — evaluation runs through the entire project lifecycle
- Training Recipes and Hyperparameter Tuning — engineering reproducibility of experiments
References
- Bergstra, Bengio. Random Search for Hyper-Parameter Optimization (2012)
- Heusel et al. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium (2017, introduced FID)
- Salimans et al. Improved Techniques for Training GANs (2016, introduced IS)
- Sokolova, Lapalme. A systematic analysis of performance measures for classification tasks (2009)
- scikit-learn metrics documentation