Skip to content

Common Pitfalls and Anti-patterns

Quick overview The most common ways things go wrong in deep learning practice, with symptoms, causes, and fixes for each: data leakage, evaluation leakage, wrong learning rate, forgetting normalization, class imbalance, unfixed seeds, validation overfitting, BatchNorm traps, mixed precision bugs, fine-tuning learning rate too high, OOM, and the illusion of "looking correct."

Common Pitfalls and Anti-patterns ​

One-sentence definition: 80% of failures in deep learning projects aren't due to "insufficient theory" — they're from repeatedly stepping into the same engineering pitfalls. This page dissects each one, giving symptoms, causes, and fixes, so that once you've encountered them, you can recognize them.

These pitfalls span the entire pipeline — data, training, evaluation, and deployment. For the complete troubleshooting methodology, see Debugging and Diagnosis; for the systematic principles that prevent them, see DL Design Principles.

1. Data Leakage: Test-set Information Leaks into Training ​

  • Symptom: offline metrics are astonishingly high (99%+), but online performance drops significantly; or the model "overfits too perfectly" on the training set.
  • Cause: statistics, feature engineering (e.g., the mean/variance used for standardization), and class encodings that should only be computed from the training set have been computed from the full dataset (including the test set).
  • Fix:
    python
    # Wrong: compute stats from the entire dataset
    mean, std = all_data.mean(), all_data.std()
    
    # Correct: fit only on training set, apply same stats to test set
    mean, std = train_data.mean(), train_data.std()
    For time-series data, watch out for temporal leakage (using future information to predict the past); for user data, watch out for identity leakage (samples from the same user must not appear in both train and test). Full checklist in Evaluation Practices.

2. Evaluation Leakage: Tuning on the Test Set ​

  • Symptom: test-set accuracy keeps "climbing" to absurd levels, but the model collapses when data is changed.
  • Cause: the test set is being used as a validation set — checking the test-set score after every experiment essentially "trains" the test set into model selection.
  • Fix: strictly enforce a train / val / test three-way split. Hyperparameter tuning, model selection, early stopping, and checkpoint selection all happen on the validation set. The test set is touched only once, for final evaluation. If you must check the test set multiple times, log every result and acknowledge that the metric is already corrupted.

3. Wrong Learning Rate: Too Large, Too Small, or No Scheduling ​

SymptomCauseFix
Loss is NaN/diverges from step oneLearning rate too largeRange testing to find the right interval (see Training Recipes and Hyperparameter Tuning)
Loss decreases extremely slowly, like "crawling"Learning rate too smallIncrease 10× and retry; if still stuck, check data/model
Training set fits, but validation stallsNo scheduling, late-stage oscillationAdd cosine annealing or halve on validation plateau
Adam doesn't convergeForgot Adam also needs LR tuning (default 1e-3 is a starting point, not the endpoint)Range testing + scheduling

The learning rate is the "cheapest, most impactful" hyperparameter — check it first when things go wrong.

4. Forgetting Normalization ​

  • Symptom: loss is higher, convergence is slower or oscillates under the same config; results vary wildly across different initializations.
  • Cause: input ranges are off (0–255 mixed with 0–1), and the scale of weight initialization doesn't match the scale of inputs (mechanics in Initialization and Normalization).
  • Fix: check x.min() / x.max() before training, normalize uniformly to [0,1] or standardize; when using pre-trained models for image tasks, you must use the mean and std the model was trained with (e.g., ImageNet's (0.485, 0.456, 0.406)).

5. Improper Handling of Class Imbalance ​

  • Symptom: accuracy looks fine (90%+), but the minority class is all zeros; F1 is abysmal.
  • Cause: default cross-entropy loss is dominated by the majority class, so the model learns to "just predict the majority class."
  • Fixes (sorted by cost-effectiveness):
    1. Change evaluation metric: don't use accuracy; use per-class precision/recall/F1 (see Evaluation Practices).
    2. Weighted loss: CrossEntropyLoss(weight=class_weights), with weights inversely proportional to class frequency.
    3. Resampling: oversample the minority class / undersample the majority class (don't do this on the test set).
    4. Change loss: Focal Loss (hard-sample weighting) for detection tasks with extreme imbalance.
    python
    # Weighted loss example
    counts = torch.bincount(train_labels, minlength=n_classes).float()
    class_weights = 1.0 / counts
    class_weights = class_weights / class_weights.sum() * n_classes
    criterion = nn.CrossEntropyLoss(weight=class_weights.to(device))

6. Unfixed Seeds Lead to Unreproducible Results ​

  • Symptom: the same code produces different results across runs; "90% yesterday, 87% today"; can't explain scores to anyone — including yourself.
  • Cause: random initialization, data shuffling, augmentation sampling, cuDNN algorithm selection — all introduce randomness, and none are being fixed.
  • Fix: fix torch/numpy/random seeds + cudnn.deterministic (code template in Training Recipes and Hyperparameter Tuning). Note: fixing seeds makes experiments reproducible, it doesn't make the model better — for evaluation, it's actually better to average over multiple seeds (see Evaluation Practices).

7. Overfitting on the Validation Set ​

  • Symptom: validation score keeps climbing until it hits 99%, then collapses immediately when data is changed.
  • Cause: the validation set is too small or viewed too many times — selecting the best checkpoint at every epoch means the validation set itself gets memorized.
  • Fix:
    • The validation set should be large enough and fixed (at least several thousand samples).
    • For early stopping and checkpoint selection, don't keep cherry-picking the best on the same validation samples; keep at least one more "isolated" final test set.
    • When making high-variance decisions (e.g., comparing two similar configs) using a small validation set, use confidence intervals to determine whether differences are significant.

8. BatchNorm Traps ​

Two high-frequency failure modes:

1. Batch size too small

  • Symptom: training loss oscillates, validation metrics are poor and unstable; batch-level statistics are noisy.
  • Cause: BatchNorm uses the current batch's mean and variance; when the batch is too small (<16), estimates are inaccurate; at inference, it uses running statistics instead, creating a training/inference behavior split.
  • Fix: keep batch ≥32 when possible; for small-batch scenarios, switch to GroupNorm/LayerNorm, or increase batch size (gradient accumulation helps).

2. Training/inference behavior mismatch

  • Symptom: metrics look fine on training data, but there's a systematic bias at test time.
  • Cause: forgetting model.eval() during evaluation, so BatchNorm is still using batch statistics instead of running statistics.
  • Fix: call model.eval() before evaluation; confirm model.train() before backpropagation.
python
# Evaluation must be written this way
model.eval()
with torch.no_grad():
    ...  # inference

9. Mixed Precision (AMP) Bugs ​

  • Symptom: intermittent NaN mid-training; some layers output all zeros; CPU and GPU results disagree.
  • Cause: float16 has a small dynamic range (upper bound ~65504); large activations or small gradients overflow/underflow; custom operators that lack fp16 kernels get silently downgraded or fail.
  • Fix:
    • Use the official GradScaler instead of hand-written scaling;
    • For custom losses/operators, check their fp16 behavior, and force fp32 for problematic parts when necessary;
    • When NaN appears, reproduce with pure fp32 first to confirm the issue is AMP-related;
    • Eliminate one by one: use torch.autocast(enabled=False) for a quick comparison.

10. Fine-tuning Learning Rate Too High ​

  • Symptom: fine-tuning loss spikes dramatically in the early stage and never recovers, or accuracy is worse than when the backbone is frozen.
  • Cause: pre-trained weights are already close to a good solution; using a default training learning rate (e.g., 1e-3) will "wash away" the learned representations.
  • Fix: fine-tuning learning rates should be 1–2 orders of magnitude smaller than training from scratch (1e-4 to 1e-5); first freeze the backbone and train only the classification head, then unfreeze for fine-tuning (recipe in Training Recipes and Hyperparameter Tuning, example in Progressive Tutorial: Three Versions).

11. Handling VRAM OOM ​

  • Symptom: CUDA out of memory; or VRAM runs out mid-training.
  • Cause: model parameters + intermediate activations + optimizer states exceed VRAM; or VRAM leak (computation graphs accumulate every batch).
  • Fixes (sorted by cost-effectiveness):
    1. Reduce batch (and adjust learning rate accordingly) or gradient accumulation.
    2. Mixed precision: saves 30–50% VRAM immediately.
    3. Gradient checkpointing: trade space for time.
    4. Check for VRAM leaks: are you calling zero_grad() before loss.backward()? Are there Tensor references held outside of .item() calls in the loop?
    5. Use torch.cuda.max_memory_allocated() to locate the peak source.

Don't ignore leaks

"OOM but reducing the batch doesn't help" is usually a VRAM leak — something is keeping a computation graph or Tensor reference alive inside the loop. Use max_memory_allocated to observe whether memory grows monotonically with each batch to confirm.

12. The Illusion of "Looking Correct" ​

  • Symptom: metrics pass, demos work, but the model collapses when data or scenarios change; or the demo is "cheating."
  • Cause: flawed evaluation design hides real problems — commonly:
    • Train-test same distribution: test set comes from the same batch as training, masking distribution shifts.
    • Data leakage (see #1) inflates metrics.
    • Tuning on the test set (see #2).
    • Only looking at averages: average metrics hide tail-end catastrophes from "good segments."
    • Demo-specific data: demo data is hand-picked, always samples the model excels at.
  • Fix: build out-of-distribution test sets (data from different time periods / regions / demographics); bucket metrics by difficulty/confidence; treat "a model can only work on its own test distribution" as the default assumption and actively verify where it fails. Interpretability tools can further reveal what the model is actually "relying on" to get things right — see Interpretability and Fairness.

Appendix: Anti-pattern Quick Reference ​

Anti-patternOne-line SignatureCore Fix
Data leakageMetrics unrealistically goodCompute stats from training set only
Evaluation leakageTest set peeked at repeatedlyThree-way split, test touched once
Wrong learning rateLoss won't go down or jumps aroundRange testing
Forgetting normalizationSlow, unstable convergenceCheck input ranges
Class imbalanceHigh accuracy but minority class failsWeighted loss / change metric
Unfixed seedsUnreproducible resultsFix seeds across the full pipeline
Validation overfitting99% on validation, collapses elsewhereFixed large validation set + isolated test set
BatchNorm trapsSmall-batch oscillation / train-inference mismatchbatch ≥32, correct eval()
AMP bugIntermittent NaNGradScaler + pure fp32 comparison
Fine-tuning lr too highFine-tuning ruins the modelLR 1–2 orders of magnitude smaller
VRAM OOMCrashes repeatedlyFour-piece set: batch/AMP/checkpoint/check for leaks
IllusionLooks fine but isn'tOOD tests + bucketed metrics

Further Reading ​

References ​