Theme
Overfitting and Regularization
One-line definition: Overfitting is when a model has "memorized the training set" rather than "learned the patterns," and regularization is the collective term for all techniques that suppress this behavior and improve generalization. How overfitting is observed and measured is covered in Deep Learning Evaluation and Experiments through learning curves; this article focuses on "what to do after diagnosis."
I. Overfitting Mechanism and Bias-Variance Decomposition
Why do models overfit? Looking at it from the framework of statistical learning, the generalization error can be decomposed into three parts:
Generalization error = bias² + variance + irreducible noise- Bias: the systematic gap between the model's hypothesis and the true function — the model is too simple (underfitting).
- Variance: how sensitive the model is to "a different batch of training data" — the model is too complex and memorizes noise (overfitting).
- Noise: the inherent randomness in the data that no one can learn.
Overfitting = excessive variance: swap the training set, and the learned function changes dramatically. The goal of machine learning is not to minimize training error, but to minimize the sum of bias and variance — this is the statistical essence of the train/validation curve divergence in Deep Learning Evaluation and Experiments.
Deep networks have a counterintuitive phenomenon: models with far more parameters than samples can still generalize well (e.g., GPT-series). Research on this mystery ("double descent," implicit regularization, etc.) is ongoing, but in practice we rely on a well-defined set of tools — each covered below.
II. L1/L2 Regularization and Weight Decay
Add a penalty term for parameter magnitude to the loss function:
L_total = L_data + λ·R(θ)- L2 regularization:
R = Σ θᵢ². Pushes weights toward 0 but not all the way — large weights are suppressed, improving the "smoothness" of the solution. In gradient descent, it is equivalent to multiplying weights by(1−ηλ)at each step, which is why it is also called weight decay. Note the discussion of "decoupled weight decay" in AdamW, see Optimization and Gradient Descent. - L1 regularization:
R = Σ |θᵢ|. Forces many weights exactly to 0, producing sparse solutions.
Why is L1 sparse while L2 is not? A geometric intuition: the contour of L1 is a diamond (with sharp corners on the axes), while L2 is a sphere. When the loss surface is tangent to these contours, the diamond tends to cut at an axis (some dimensions become 0), while the sphere typically cuts at a non-axis position (all small but nonzero). L1 is therefore commonly used for feature selection and interpretable models; L2 is the universal default.
III. Dropout and Variants
Dropout (Hinton et al., 2012): during training, each neuron is randomly zeroed with probability p (output multiplied by 1/(1−p) to keep the expectation unchanged); during inference, all neurons are kept. Mechanistic explanations:
- Ensemble effect: each training step trains a different "subnetwork," and inference is equivalent to an approximate average of exponentially many subnetworks (bagging).
- Preventing co-adaptation: neurons can't rely on the existence of specific "partners," forcing them to learn more independent features — this is the core problem that "co-adaptation" prevention addresses.
Variant family:
- SpatialDropout: zero out entire channels (not individual pixels), commonly used in CNNs and embedding layers, disrupting "channel-level" co-adaptation.
- DropBlock (2018): contiguous block-wise zeroing, regularizes convolutional layers more effectively than pointwise Dropout.
- DropConnect: zero out weights rather than activations — stronger but more computationally expensive.
- Stochastic Depth: randomly skip entire residual blocks during training — turning "deep" into "an ensemble of shallow networks," with effects similar to a depth version of Dropout.
PyTorch: nn.Dropout(p), nn.Dropout2d (Spatial). Note that Dropout is automatically disabled under model.eval() — like BN, this is another example of different train/inference behavior (see Initialization and Normalization).
IV. Early Stopping
Truncate training at the point where validation loss stops decreasing (or starts rising). It is essentially an implicit regularization that minimizes "how far parameters are from the starting point" — parameters stay in a small-norm region closer to initialization, corresponding to a smoother solution. Implementation details:
- On the validation set, "best" is judged by no improvement for N consecutive epochs (patience), to guard against noise-induced false signals.
- Save the best weights (
torch.save(model.state_dict())) and restore them at the end of training. - Early stopping is "free regularization" — every project should have it. The validation set is used repeatedly by early stopping, so the final score is always settled on the test set (see the three-set split discipline in the Evaluation section).
V. Data Augmentation
Create more, more realistic training samples through transformations — the most effective and "honest" form of regularization — because it teaches the model invariance rather than suppressing it.
- Images: random crop/flip/rotate, color jittering, zoom, translate, blur. Strong augmentations (AutoAugment, RandAugment, Cutout) are standard for ImageNet SOTA.
- Text: synonym replacement, back-translation, random word deletion/reordering, EDA methods. Note: text augmentation must preserve semantics; destroying grammar introduces noise.
- Audio: time stretching, pitch/volume perturbation, SpecAugment (time/frequency masking). See Speech and Audio.
Judging augmentation intensity: too weak is basically doing nothing, too strong turns "intra-class variation" into "cross-class confusion." Regulate based on validation set performance. The engineering of data pipelines (online augmentation, multi-GPU shuffling) is covered in Data and Data Engineering.
VI. Label Smoothing
Replace one-hot hard labels with soft labels: y' = (1−ε)·y + ε/K (ε≈0.1). Its dual role:
- Suppresses overfitting: the model is no longer forced to push the correct class probability to 1, so logits don't grow without bound.
- Improves calibration: predicted probabilities more closely match true confidence. See Loss Functions and Output Layers.
Side effect: soft labels reduce the benefit of "distillation"-type methods (the teacher model's predictions are flattened). It is a "cheap, nearly side-effect-free" regularizer and is the default for the Transformer family.
VII. EMA (Exponential Moving Average)
Maintain a running average copy of the parameters:
θ_ema ← β·θ_ema + (1−β)·θ # β≈0.999Use θ_ema for inference instead of the current θ. Why it works: parameters jitter at high frequency during training, and EMA is equivalent to "averaging over the time dimension on the parameter trajectory," denoising into flatter, better-generalizing regions. It is virtually mandatory in GANs, diffusion models, and self-supervised learning — for example, in VAEs and GANs, generation quality is noticeably more stable using EMA weights after discriminator/generator training. In implementation, PyTorch offers torch.optim.swa_utils.AveragedModel or the torch_ema library.
VIII. Mixup and CutMix
Mixup (2018): linearly interpolate two samples and their labels:
x' = λ·xᵢ + (1−λ)·xⱼ, y' = λ·yᵢ + (1−λ)·yⱼThe model learns "linear interpolation between categories" — training boundaries become smoother and more robust to adversarial examples. CutMix (2019): cut and paste an image patch from one sample onto another, mixing labels by area ratio — combining the smoothness of Mixup with the preservation of local features.
These "mixed augmentations" are a free lunch in CV (stable performance gains), and have variants for text like Manifold Mixup (interpolation in the hidden space). Note that they can be too aggressive when stacked with strong augmentation; use the validation set as a gate.
IX. Regularization Combination Strategies and Intensity Selection
There are many regularization techniques, but total regularization intensity is what matters — all techniques push in the same direction: reducing variance. Combination experience:
- Deploy from "least destructive" to "most destructive": early stopping + data augmentation → L2/weight decay → Dropout → stronger augmentation / reduce capacity. Add cheap methods first; don't stack everything from the start.
- The presence of normalization layers changes the need: BN has a mild regularizing effect (see Initialization and Normalization), so CNNs using BN often see very little benefit from Dropout.
- Intensity is set by the validation set: select regularization hyperparameters (λ, p, ε) on the validation set — don't guess. The validation set is a "repeatedly usable" selection set; don't treat it as a test set.
- Start large, then tighten: default to selecting a slightly oversized network + sufficient regularization is easier to tune than a "small network trained hard" — regularization gives capacity room to make mistakes.
Don't Over-regularize
Regularization is essentially "a bias toward simpler solutions." If the model is underfitting (can't even bring down the training loss), adding regularization only makes things worse — first confirm it is overfitting (good on training, bad on validation), and reading the learning curve is the first step.
X. Tradeoffs
Tradeoffs
Bias vs. variance: this is the master framework for all regularization. When data volume is sufficient, the returns from regularization diminish and can even be harmful (suppressing things the model should have learned); when data is scarce, regularization is a lifesaver. Data volume determines regularization intensity.
L1's sparsity and interpretability vs. L2's smoothness and stability: L1 is suitable for feature selection and compression; L2 is universal and optimizes more smoothly. Deep networks default to L2/weight decay; L1 is mostly used for sparsification (e.g., as a preprocessor for pruning).
Dropout's cost vs. benefit: Dropout amplifies activations (multiplying by 1/(1−p)), which costs some capacity. For very large models, the impact is negligible; for small models, it might be "too much regularization." Remember the eval() mode at inference time.
The benefit of data augmentation vs. computational cost: online augmentation increases training time; its "cost-effectiveness" for improving generalization is almost always the highest — but strong augmentations can distort distributions (e.g., flipping is invalid for medical images), so it must be combined with task semantics.
Regularization is not a "fancy ornament" — it determines whether a model can upgrade from "memorizing the training set" to "mastering the task." It, together with Optimization and Gradient Descent (optimization itself also carries implicit regularization) and the mild regularization of normalization layers, forms the "generalization triangle." The full flow for combinatorial tuning is covered in Training Recipes and Hyperparameter Tuning; terminology is in Glossary.
Further Reading
- Neural Networks Fundamentals — the relationship between capacity and overfitting
- Data and Data Engineering — the engineering of augmentation and data quality
- Loss Functions and Output Layers — implementation details of label smoothing
- Training Recipes and Hyperparameter Tuning — the full tuning flow for regularization intensity
- Common Pitfalls and Anti-patterns — common errors related to regularization
- Debugging and Diagnostics — quick localization of overfitting symptoms
References
- Srivastava et al. Dropout: A Simple Way to Prevent Neural Networks from Overfitting (2014)
- Zhang et al. mixup: Beyond Empirical Risk Minimization (2018)
- Yun et al. CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features (2019)
- Ghiasi et al. DropBlock: A regularization method for convolutional networks (2018)
- Szegedy et al. Rethinking the Inception Architecture for Computer Vision (2016, Label Smoothing)
- Huang et al. Deep Networks with Stochastic Depth (2016)