Theme
Common Pitfalls and Anti-patterns
One-sentence definition: 80% of failures in deep learning projects aren't due to "insufficient theory" — they're from repeatedly stepping into the same engineering pitfalls. This page dissects each one, giving symptoms, causes, and fixes, so that once you've encountered them, you can recognize them.
These pitfalls span the entire pipeline — data, training, evaluation, and deployment. For the complete troubleshooting methodology, see Debugging and Diagnosis; for the systematic principles that prevent them, see DL Design Principles.
1. Data Leakage: Test-set Information Leaks into Training
- Symptom: offline metrics are astonishingly high (99%+), but online performance drops significantly; or the model "overfits too perfectly" on the training set.
- Cause: statistics, feature engineering (e.g., the mean/variance used for standardization), and class encodings that should only be computed from the training set have been computed from the full dataset (including the test set).
- Fix:pythonFor time-series data, watch out for temporal leakage (using future information to predict the past); for user data, watch out for identity leakage (samples from the same user must not appear in both train and test). Full checklist in Evaluation Practices.
# Wrong: compute stats from the entire dataset mean, std = all_data.mean(), all_data.std() # Correct: fit only on training set, apply same stats to test set mean, std = train_data.mean(), train_data.std()
2. Evaluation Leakage: Tuning on the Test Set
- Symptom: test-set accuracy keeps "climbing" to absurd levels, but the model collapses when data is changed.
- Cause: the test set is being used as a validation set — checking the test-set score after every experiment essentially "trains" the test set into model selection.
- Fix: strictly enforce a train / val / test three-way split. Hyperparameter tuning, model selection, early stopping, and checkpoint selection all happen on the validation set. The test set is touched only once, for final evaluation. If you must check the test set multiple times, log every result and acknowledge that the metric is already corrupted.
3. Wrong Learning Rate: Too Large, Too Small, or No Scheduling
| Symptom | Cause | Fix |
|---|---|---|
| Loss is NaN/diverges from step one | Learning rate too large | Range testing to find the right interval (see Training Recipes and Hyperparameter Tuning) |
| Loss decreases extremely slowly, like "crawling" | Learning rate too small | Increase 10× and retry; if still stuck, check data/model |
| Training set fits, but validation stalls | No scheduling, late-stage oscillation | Add cosine annealing or halve on validation plateau |
| Adam doesn't converge | Forgot Adam also needs LR tuning (default 1e-3 is a starting point, not the endpoint) | Range testing + scheduling |
The learning rate is the "cheapest, most impactful" hyperparameter — check it first when things go wrong.
4. Forgetting Normalization
- Symptom: loss is higher, convergence is slower or oscillates under the same config; results vary wildly across different initializations.
- Cause: input ranges are off (0–255 mixed with 0–1), and the scale of weight initialization doesn't match the scale of inputs (mechanics in Initialization and Normalization).
- Fix: check
x.min()/x.max()before training, normalize uniformly to [0,1] or standardize; when using pre-trained models for image tasks, you must use the mean and std the model was trained with (e.g., ImageNet's(0.485, 0.456, 0.406)).
5. Improper Handling of Class Imbalance
- Symptom: accuracy looks fine (90%+), but the minority class is all zeros; F1 is abysmal.
- Cause: default cross-entropy loss is dominated by the majority class, so the model learns to "just predict the majority class."
- Fixes (sorted by cost-effectiveness):
- Change evaluation metric: don't use accuracy; use per-class precision/recall/F1 (see Evaluation Practices).
- Weighted loss:
CrossEntropyLoss(weight=class_weights), with weights inversely proportional to class frequency. - Resampling: oversample the minority class / undersample the majority class (don't do this on the test set).
- Change loss: Focal Loss (hard-sample weighting) for detection tasks with extreme imbalance.
python# Weighted loss example counts = torch.bincount(train_labels, minlength=n_classes).float() class_weights = 1.0 / counts class_weights = class_weights / class_weights.sum() * n_classes criterion = nn.CrossEntropyLoss(weight=class_weights.to(device))
6. Unfixed Seeds Lead to Unreproducible Results
- Symptom: the same code produces different results across runs; "90% yesterday, 87% today"; can't explain scores to anyone — including yourself.
- Cause: random initialization, data shuffling, augmentation sampling, cuDNN algorithm selection — all introduce randomness, and none are being fixed.
- Fix: fix torch/numpy/random seeds +
cudnn.deterministic(code template in Training Recipes and Hyperparameter Tuning). Note: fixing seeds makes experiments reproducible, it doesn't make the model better — for evaluation, it's actually better to average over multiple seeds (see Evaluation Practices).
7. Overfitting on the Validation Set
- Symptom: validation score keeps climbing until it hits 99%, then collapses immediately when data is changed.
- Cause: the validation set is too small or viewed too many times — selecting the best checkpoint at every epoch means the validation set itself gets memorized.
- Fix:
- The validation set should be large enough and fixed (at least several thousand samples).
- For early stopping and checkpoint selection, don't keep cherry-picking the best on the same validation samples; keep at least one more "isolated" final test set.
- When making high-variance decisions (e.g., comparing two similar configs) using a small validation set, use confidence intervals to determine whether differences are significant.
8. BatchNorm Traps
Two high-frequency failure modes:
1. Batch size too small
- Symptom: training loss oscillates, validation metrics are poor and unstable; batch-level statistics are noisy.
- Cause: BatchNorm uses the current batch's mean and variance; when the batch is too small (<16), estimates are inaccurate; at inference, it uses running statistics instead, creating a training/inference behavior split.
- Fix: keep batch ≥32 when possible; for small-batch scenarios, switch to
GroupNorm/LayerNorm, or increase batch size (gradient accumulation helps).
2. Training/inference behavior mismatch
- Symptom: metrics look fine on training data, but there's a systematic bias at test time.
- Cause: forgetting
model.eval()during evaluation, so BatchNorm is still using batch statistics instead of running statistics. - Fix: call
model.eval()before evaluation; confirmmodel.train()before backpropagation.
python
# Evaluation must be written this way
model.eval()
with torch.no_grad():
... # inference9. Mixed Precision (AMP) Bugs
- Symptom: intermittent NaN mid-training; some layers output all zeros; CPU and GPU results disagree.
- Cause: float16 has a small dynamic range (upper bound ~65504); large activations or small gradients overflow/underflow; custom operators that lack fp16 kernels get silently downgraded or fail.
- Fix:
- Use the official
GradScalerinstead of hand-written scaling; - For custom losses/operators, check their fp16 behavior, and force fp32 for problematic parts when necessary;
- When NaN appears, reproduce with pure fp32 first to confirm the issue is AMP-related;
- Eliminate one by one: use
torch.autocast(enabled=False)for a quick comparison.
- Use the official
10. Fine-tuning Learning Rate Too High
- Symptom: fine-tuning loss spikes dramatically in the early stage and never recovers, or accuracy is worse than when the backbone is frozen.
- Cause: pre-trained weights are already close to a good solution; using a default training learning rate (e.g.,
1e-3) will "wash away" the learned representations. - Fix: fine-tuning learning rates should be 1–2 orders of magnitude smaller than training from scratch (
1e-4to1e-5); first freeze the backbone and train only the classification head, then unfreeze for fine-tuning (recipe in Training Recipes and Hyperparameter Tuning, example in Progressive Tutorial: Three Versions).
11. Handling VRAM OOM
- Symptom:
CUDA out of memory; or VRAM runs out mid-training. - Cause: model parameters + intermediate activations + optimizer states exceed VRAM; or VRAM leak (computation graphs accumulate every batch).
- Fixes (sorted by cost-effectiveness):
- Reduce batch (and adjust learning rate accordingly) or gradient accumulation.
- Mixed precision: saves 30–50% VRAM immediately.
- Gradient checkpointing: trade space for time.
- Check for VRAM leaks: are you calling
zero_grad()beforeloss.backward()? Are there Tensor references held outside of.item()calls in the loop? - Use
torch.cuda.max_memory_allocated()to locate the peak source.
Don't ignore leaks
"OOM but reducing the batch doesn't help" is usually a VRAM leak — something is keeping a computation graph or Tensor reference alive inside the loop. Use max_memory_allocated to observe whether memory grows monotonically with each batch to confirm.
12. The Illusion of "Looking Correct"
- Symptom: metrics pass, demos work, but the model collapses when data or scenarios change; or the demo is "cheating."
- Cause: flawed evaluation design hides real problems — commonly:
- Train-test same distribution: test set comes from the same batch as training, masking distribution shifts.
- Data leakage (see #1) inflates metrics.
- Tuning on the test set (see #2).
- Only looking at averages: average metrics hide tail-end catastrophes from "good segments."
- Demo-specific data: demo data is hand-picked, always samples the model excels at.
- Fix: build out-of-distribution test sets (data from different time periods / regions / demographics); bucket metrics by difficulty/confidence; treat "a model can only work on its own test distribution" as the default assumption and actively verify where it fails. Interpretability tools can further reveal what the model is actually "relying on" to get things right — see Interpretability and Fairness.
Appendix: Anti-pattern Quick Reference
| Anti-pattern | One-line Signature | Core Fix |
|---|---|---|
| Data leakage | Metrics unrealistically good | Compute stats from training set only |
| Evaluation leakage | Test set peeked at repeatedly | Three-way split, test touched once |
| Wrong learning rate | Loss won't go down or jumps around | Range testing |
| Forgetting normalization | Slow, unstable convergence | Check input ranges |
| Class imbalance | High accuracy but minority class fails | Weighted loss / change metric |
| Unfixed seeds | Unreproducible results | Fix seeds across the full pipeline |
| Validation overfitting | 99% on validation, collapses elsewhere | Fixed large validation set + isolated test set |
| BatchNorm traps | Small-batch oscillation / train-inference mismatch | batch ≥32, correct eval() |
| AMP bug | Intermittent NaN | GradScaler + pure fp32 comparison |
| Fine-tuning lr too high | Fine-tuning ruins the model | LR 1–2 orders of magnitude smaller |
| VRAM OOM | Crashes repeatedly | Four-piece set: batch/AMP/checkpoint/check for leaks |
| Illusion | Looks fine but isn't | OOD tests + bucketed metrics |
Further Reading
- Debugging and Diagnosis — systematic troubleshooting for the pitfalls on this page
- DL Design Principles — experimental discipline that prevents pitfalls
- Training Recipes and Hyperparameter Tuning — correct approaches for learning rate, seeds, and fine-tuning
- Evaluation Practices — the antidote to evaluation leakage and "looking correct"
- Build Your Own DL Project — where most of these pitfalls first appear
- Initialization and Normalization — theoretical roots of normalization-related pitfalls
References
- Karpathy. A Recipe for Training Neural Networks — classic discussion of the "looks effective but isn't" trap
- Kaufman et al. Leakage and the Reproducibility Crisis in ML-based Science (2023) — systematic study of data leakage
- Paszke et al. Automatic differentiation in PyTorch (2017) — mechanistic background on AMP and autograd