Theme
Training Recipes and Hyperparameter Tuning
In one sentence: A training recipe is a set of proven default choices — data, initialization, loss, optimizer, learning rate, regularization — combined into a solid starting point for reproducible convergence; tuning is searching along one variable at a time from that starting point, using measurable signals (loss, validation metrics).
Tuning isn't alchemy — it's an engineering discipline with defaults, diagnostic signals, and search methods. Theoretical foundations (gradient descent, loss functions, regularization) are in Optimization and Gradient Descent and Overfitting and Regularization; troubleshooting steps for "loss not decreasing" and similar symptoms are in Debugging and Diagnostics.
1. Default Recipe Table
When in doubt, copy this table first. It covers a reasonable starting point for most vision and text tasks:
| Component | Default choice | Notes |
|---|---|---|
| Data normalization | Mean/std standardization (computed on training set only) | Inputs near zero mean, unit variance |
| Data augmentation | Vision: random crop + flip; Text: random mask/crop | Augmentation intensity increases with data size |
| Weight initialization | PyTorch default (Kaiming init + correct fan mode) | Manual initialization is a common beginner mistake |
| Loss function | Classification: CrossEntropyLoss; Regression: MSELoss/HuberLoss | Must match the output layer (see Loss Functions and Output Layers) |
| Optimizer | Default: Adam (lr=1e-3); CNNs: SGD+momentum (lr=1e-2 with weight_decay) | Adam is insensitive to lr and easy to get started with |
| Learning rate | 3e-4 (Adam) ~ 1e-2 (SGD), then refine with range test | See next section |
| Batch size | 32–256, as large as possible (powers of 2) | Interacts with learning rate, see "batch size and learning rate" |
| Normalization layer | BatchNorm (CV); LayerNorm (NLP/Transformer) | See Initialization and Normalization |
| Regularization | weight_decay (Adam with 1e-4~5e-4); Dropout as needed | Increase when data is small / model is large |
| Epochs | Small data: 10–50; Large data: 50–200, with early stopping | Early stopping on validation set is the final arbiter |
| Learning rate scheduling | Warmup + cosine annealing (or halve on validation plateau) | See "training schedule planning" |
| Random seed | torch.manual_seed + np.random.seed + env variables | See "experiment reproducibility discipline" |
One-line principle
Get training to "overfit normally" before talking about regularization. If a model can't even fit the training set (training loss doesn't decrease), the problem is in the model or data, not regularization; adding weight_decay at that point makes things worse.
2. How to Pick Learning Rate: Range Test
Learning rate is the most important single hyperparameter. Too small → slow convergence; too large → divergence. Rather than guessing, do a learning rate range test: sweep the learning rate from very small to relatively large over one epoch (linearly or exponentially), and plot the "learning rate → loss" curve:
python
import torch, torch.nn as nn
from torch.utils.data import DataLoader
def lr_find(model, loader, criterion, optimizer_cls,
lr_min=1e-6, lr_max=1e0, steps=200):
"""Sweep learning rate across `steps` batches, return (lrs, losses)"""
optimizer = optimizer_cls(model.parameters(), lr=lr_min)
scheduler = torch.optim.lr_scheduler.LambdaLR(
optimizer, lambda t: lr_min * (lr_max / lr_min) ** (t / steps))
lrs, losses = [], []
for i, (x, y) in enumerate(loader):
if i >= steps:
break
optimizer.zero_grad()
loss = criterion(model(x), y)
loss.backward()
optimizer.step()
scheduler.step()
lrs.append(optimizer.param_groups[0]["lr"])
losses.append(loss.item())
return lrs, lossesLook at the plot: take the learning rate at roughly 1/10 of the steepest descent region where the loss first starts dropping. For example, if loss starts dropping noticeably at 1e-4 and begins rebounding at 1e-2, pick something around 3e-3.
The shape of this curve is itself diagnostic information: if no learning rate makes the loss decrease, the problem is likely not the learning rate but the data or model (troubleshooting steps in Debugging and Diagnostics). PyTorch also provides built-in options via torch.optim.lr_scheduler. See the original paper by Leslie Smith in the references below.
3. Batch Size and Learning Rate Interplay
Batch size is not an isolated hyperparameter — it couples with learning rate and gradient quality:
- Larger batch gradients are more accurate (lower variance). You can and should use larger learning rates. Rule of thumb: doubling the batch size means multiplying the learning rate by roughly √2 (the square-root version of the linear scaling rule).
- Smaller batch gradients have more noise, which provides implicit regularization but is less stable to converge; typically pair with a smaller learning rate.
- Large batch final generalization can sometimes be slightly worse than small batch ("large-batch generalization gap"), but the prerequisite is that you also scale the learning rate and scheduling proportionally — in many cases the gap comes from the learning rate not being tuned, not from the batch size itself.
python
# Example of complementary changes when moving from batch_size 32 to 64
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3 * (64 / 32) ** 0.5)Low VRAM is not an excuse for small batch size
When VRAM limits batch size, prioritize gradient accumulation (see "memory management") to simulate a large batch rather than directly reducing batch size while keeping the original learning rate — the latter silently changes the training dynamics.
4. Training Schedule: Warmup + Cosine Annealing
Modern training is typically not a constant learning rate from start to finish, but rather a three-phase approach:
- Warmup (a few hundred to a few thousand steps): Learning rate increases linearly from a small value to the target. Reason: gradient statistics haven't stabilized in early training (especially with BatchNorm); jumping straight to a large learning rate can mislead the optimizer (Adam's momentum/variance estimates) toward early samples.
- Main phase: Maintain or slowly decay.
- Cosine annealing / linear decay: Smoothly bring the learning rate close to 0, letting parameters "refine" near the minimum of the loss landscape.
python
import math
from torch.optim.lr_scheduler import LambdaLR
def cosine_with_warmup(optimizer, total_steps, warmup_steps=1000):
def lr_lambda(step):
if step < warmup_steps:
return step / warmup_steps # Linear warmup
progress = (step - warmup_steps) / max(1, total_steps - warmup_steps)
return 0.5 * (1 + math.cos(math.pi * progress)) # Cosine decay to 0
return LambdaLR(optimizer, lr_lambda)
scheduler = cosine_with_warmup(optimizer, total_steps=len(train_loader) * 30)torch.optim.lr_scheduler also offers more convenient options like CosineAnnealingWarmRestarts and ReduceLROnPlateau (which directly watches the validation set). The choice of learning rate scheduling depends on the loss landscape shape — it's an empirical art — but "warmup + decay" is almost always better than constant learning rate.
5. Regularization Combination Intensity
Regularization is "scale according to need," not "more is better." The general logic of the combo punch:
| Signal | Regularization approach | Intensity recommendation |
|---|---|---|
| Low training loss, high validation loss (overfitting) | weight_decay, Dropout, augmentation, early stopping | Start with the cheapest: weight_decay, then add one at a time |
| Training loss won't go down (underfitting) | Turn off regularization; check model capacity / learning rate | Don't add more regularization |
| Very small data (<10K) | All standard regularizations + early stopping + transfer learning | Strongly recommend transfer learning |
| Very large training data (>1M) | Regularization has limited effect; focus on scheduling and data quality | Lightweight weight_decay is sufficient |
The "cheapest" order: weight_decay costs essentially nothing; Dropout only activates during training; augmentation has data loading overhead; early stopping is free but only guards the overfitting half of the curve. Quantify the net gain of each addition — compare "with/without" on the validation set rather than stacking by feel (experiment discipline in DL Design Principles).
6. Transfer Learning Recipe
Moving a pre-trained model to a new task has a mature two-phase recipe (theory in Representation Learning and Pre-training, example in Progressive Tutorial: Three Iterations):
- Phase 1: Freeze the backbone, train only the new classification head. Set lr to
1e-3(Adam), run for a few epochs to let the head align with the features. - Phase 2: Unfreeze all (or the latter half) parameters, fine-tune with small lr. Set lr to
1e-4~1e-5, one order of magnitude smaller than Phase 1 — pre-trained weights are already excellent; large lr will "blow out" the learned representations. - Normalization statistics must match the pre-training (ImageNet's
(0.485, 0.456, 0.406)). Otherwise the feature distribution is misaligned. - Resolution: Resize input to the pre-trained model's expected resolution (e.g., 224×224), or pick a pre-trained model designed for small images.
- Monitor fine-tuning magnitude: Compute per-layer relative weight changes. If any layer changes by more than ~10%, the learning rate is too large.
A common mistake is "full fine-tuning with from-scratch training lr" — this is the #1 cause of fine-tuning failures in Common Pitfalls and Anti-Patterns.
7. Memory Management Toolkit
When the model doesn't fit in VRAM, apply these in order of cost-effectiveness:
7.1 Gradient Accumulation
Simulates a larger batch without changing gradient statistics:
python
accumulation_steps = 4 # Equivalent to 4x batch_size
optimizer.zero_grad()
for i, (x, y) in enumerate(loader):
loss = criterion(model(x), y) / accumulation_steps # Average, not sum
loss.backward()
if (i + 1) % accumulation_steps == 0:
optimizer.step()
optimizer.zero_grad()Note: divide loss by accumulation_steps to average, otherwise the effective learning rate is scaled up by accumulation_steps times. BatchNorm statistics are still computed per single batch, so large-batch BN behavior can't be fully replicated.
7.2 Mixed Precision (AMP)
Using PyTorch's built-in torch.amp, a single GPU typically saves 30–50% VRAM and speeds up, with negligible precision loss on most tasks:
python
from torch.amp import GradScaler, autocast
scaler = GradScaler("cuda")
for x, y in loader:
x, y = x.to(device), y.to(device)
optimizer.zero_grad()
with autocast("cuda"):
loss = criterion(model(x), y)
scaler.scale(loss).backward()
scaler.step(optimizer)
scaler.update()Pitfalls of mixed precision (precision overflow, NaN gradients, some ops not supporting half-precision) are covered in Common Pitfalls and Anti-Patterns.
7.3 Gradient Checkpointing
torch.utils.checkpoint trades memory for computation (re-computing the forward pass during backward):
python
from torch.utils.checkpoint import checkpoint
# Enable for sub-module forward passes
def forward(self, x):
x = checkpoint(self.layer1, x) # Don't save activations in forward, recompute in backward
x = checkpoint(self.layer2, x)
return xTypical use case: fine-tuning 7B-class large models, paired with quantization/LoRA; small models with fewer than 10 layers see little benefit.
Priority order: Mixed precision (almost zero cost) → Gradient accumulation (no model changes) → Gradient checkpointing (changes the model, slower) → Model parallelism / quantization (requires refactoring). Systematic solutions to memory issues also involve torch.utils.benchmark and memory profilers.
8. Experiment Reproducibility Discipline
"It was working fine yesterday, but now it doesn't" — almost always the environment or randomness. Reproducibility discipline is the bottom line of experimentation:
- Fix seeds: Set immediately after imports, covering torch / numpy / random:
python
import random, numpy as np, torch
def set_seed(seed):
random.seed(seed)
np.random.seed(seed)
torch.manual_seed(seed)
torch.cuda.manual_seed_all(seed)
torch.backends.cudnn.deterministic = True
torch.backends.cudnn.benchmark = False # Determinism first, slight performance trade-off- Pin dependencies:
pip freeze > requirements.txt(or lock versions inpyproject.toml), noting Python/CUDA/driver versions. - Fix the environment: Use Docker or conda environments, noting the image tag.
- Record everything: For each experiment, log hyperparameters, data version, code commit, random seed, and per-run results. Configuration-driven experiments are recommended (see DL Design Principles) — write hyperparameters into
config.yamlrather than scattering them in code. - Two "golden reproductions": Before shipping code or submitting a paper, run from scratch twice with the same seed; the error should be within the natural variance of the seed. This variance itself is the upper bound on your evaluation metric's error and should be reported in Evaluation in Practice.
Further Reading
- Debugging and Diagnostics — troubleshooting flow when a recipe doesn't work
- Progressive Tutorial: Three Iterations — progressive implementation of recipes
- Building a Deep Learning Project from Scratch — minimal pipeline
- Common Pitfalls and Anti-Patterns — the most common recipe failures
- DL Design Principles — meta-rules for experiment design
- Initialization and Normalization — the theory behind the recipe table
References
- Smith. Cyclical Learning Rates for Training Neural Networks (2017) — Original paper on learning rate range testing
- Goyal et al. Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour (2017) — Linear scaling rule (batch size and learning rate interplay)
- Loshchilov, Hutter. SGDR: Stochastic Gradient Descent with Warm Restarts (ICLR 2017) — Cosine annealing learning rate
- PyTorch. Automatic Mixed Precision — Official AMP documentation