Skip to content

Optimization and Gradient Descent

Quick overview Gradient descent is the engine that makes neural networks "learn." This article starts from the learning rate intuition, traces the evolution of SGD, momentum, Adam/AdamW, covers learning rate scheduling, gradient clipping, and batch size scaling rules, and finally presents a ready-to-use engineering defaults table.

Optimization and Gradient Descent ​

One-line definition: Optimization is the process of "stepping the loss as low as possible along the gradient" — optimizers in deep learning answer the question of "which direction and how far do parameters move next?" How gradients are computed is handled by Backpropagation and Automatic Differentiation; how to use gradients is the topic of this article.

I. Gradient Descent Principle and Learning Rate Intuition ​

The loss function L(θ) is a "terrain map" in parameter space. The gradient ∇L(θ) points in the direction of steepest ascent, so parameters move downhill along the negative gradient:

θ ← θ − η · ∇L(θ)

η (learning rate) is the step size, and it is the most sensitive of all hyperparameters:

  • η too large: overshoots valleys, oscillates back and forth, or even diverges (loss → NaN).
  • η too small: moves too slowly, making little progress after a long time and wasting compute.
  • η just right: stable descent, converges.

A rule of thumb: the learning rate determines whether training is "steady or aggressive," and momentum plus scheduling determine "how it transitions from aggressive to steady." All the optimizers that follow are essentially trying to answer "what step size should we give?"

II. SGD and Stochasticity ​

Naive gradient descent uses the average gradient of all data, making each update prohibitively expensive. Stochastic Gradient Descent (SGD) instead: compute the gradient on a small random subset (mini-batch) and update.

Stochasticity is not a bug but a feature:

  1. Saves compute: the gradient of a mini-batch is an unbiased estimate of the full gradient, and noise cancels out when the sample is large enough.
  2. Built-in implicit regularization: the noise from batch gradients can help parameters escape sharp local minima (sharp minima typically generalize poorly), which is directly related to the generalization discussion in Overfitting and Regularization.
  3. Batch size is a new degree of freedom: larger → closer to full gradient, more stable, but more expensive; smaller → more noise, more oscillation.

III. Momentum: Momentum and Nesterov ​

SGD oscillates back and forth in "long, narrow valleys," moving forward slowly. Momentum simulates physical inertia: it retains the historical gradient direction, and the current update = decayed historical direction + current gradient:

v ← γv + ∇L(θ)          # γ ≈ 0.9, historical momentum
θ ← θ − η·v

Benefit: lateral oscillations in the valley cancel out, while the forward (downhill) component keeps compounding — accelerating convergence. Nesterov momentum goes one step further: first "probe" one step in the momentum direction, then compute the gradient at that position ("look ahead before stepping"), which converges more stably and quickly. This is PyTorch's SGD(momentum=0.9, nesterov=True).

IV. Adaptive Methods: AdaGrad, RMSProp, Adam, and AdamW ​

Momentum is about "accelerating in the same direction"; another approach is to "give each parameter its own step size" — small steps for important directions, large steps for less important ones.

MethodCore IdeaKey Characteristics and Issues
AdaGrad (2011)Scales learning rate by accumulated historical gradient squaredAccumulation only increases, so the learning rate monotonically decays to 0, causing premature stopping
RMSProp (2012)Uses exponential moving average instead of accumulationLearning rate no longer monotonically goes to zero; suitable for non-stationary objectives
Adam (2015)RMSProp + momentum + bias correctionCombines the best of both, nearly requires no tuning, and has become the de facto standard
AdamW (2019)Decouples weight decay from the gradient updateSolves the coupling problem between L2 and adaptive learning rates in Adam; the standard for LLM training

Adam's update rule (with first moment m, second moment v, and bias correction):

m ← β1·m + (1−β1)·∇L        v ← β2·v + (1−β2)·∇L²
θ ← θ − η·m̂/(√v̂ + ε)

Why is AdamW better? In Adam's θ ← θ − η·λ·θ, L2 regularization is scaled by 1/√v̂, meaning parameters with small gradients get regularized more heavily, and it is coupled with η. AdamW decouples "decay" from the gradient, applying it directly to the parameter: θ ← θ·(1−ηλ) − η·m̂/√v̂ — after decoupling, the semantics of weight decay are clear and generalization is better. BERT and GPT-series all use it.

V. Learning Rate Scheduling: Warmup, Step, Cosine Annealing ​

A single, fixed learning rate throughout training is rarely optimal. Common scheduling strategies:

  • Warmup: linearly/polynomially increase from a small learning rate to the target value over the first few thousand steps. Almost mandatory for large model training — avoids pushing parameters too far with a large step size when initialization is unstable (gradient noise is high in the early stages).
  • Step decay: multiply by γ (e.g., ×0.1) every fixed number of epochs. Classic, easy to tune.
  • Cosine Annealing: the learning rate smoothly decays from the peak to near 0 along a cosine curve. Paired with warmup (warmup → peak → cosine decay) is the standard recipe for modern training. There is also "cosine with restarts" (SGDR), which uses periodic temperature increases to escape local minima.

Scheduling strategy must be paired with the optimizer: Adam-family optimizers are less sensitive to scheduling (adaptive step size), while SGD + momentum is highly dependent on it — a common experience combo is "SGD + large momentum + cosine annealing" for chasing SOTA, while "Adam/AdamW + warmup + cosine" is the stable default.

VI. Gradient Clipping ​

When the gradient norm is too large (gradient explosion; root cause in Backpropagation and Automatic Differentiation), a single update step sends parameters flying. Gradient clipping limits the norm of the gradient to a threshold:

if ‖∇L‖ > C:  ∇L ← C·∇L/‖∇L‖

In PyTorch, use torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0). It is virtually essential in RNN, Transformer, and GAN training. Note: it clips the "gradient norm," not the "loss value" — don't confuse them.

VII. Batch Size and Learning Rate Coupling: Linear Scaling ​

An empirical law (Goyal et al., 2017, large-scale ImageNet training): when batch size doubles, the learning rate should roughly also double. The reasoning: the gradient of n samples is the sum of n independent gradients, and the step size should scale up proportionally to maintain the same "effective number of steps per epoch."

But this law only holds below a critical batch size: beyond a certain threshold, the gradient of a large batch is already a good approximation of the full gradient, and further increasing the learning rate becomes unstable, requiring longer warmup or a smaller peak LR. In practice, large batch sizes are often paired with more warmup and higher learning rates.

Engineering Memory

Small batch (32–128) with AdamW + warmup + cosine annealing is the most hassle-free default; for large batch (>1024), remember to scale the learning rate proportionally and extend warmup, or generalization will typically degrade.

VIII. Intuition About Loss Surface Geometry: Saddle Points and Local Minima ​

High-dimensional loss surfaces differ from low-dimensional intuition: local minima are common in low dimensions, but true "valleys" are rare in high dimensions — saddle points are more common (some directions go downhill, others uphill). Dauphin et al. pointed out in 2014 that near saddle points, the gradient is close to 0 but it is not a minimum — naive SGD gets "stuck" there.

This is why adaptive methods (second moments independently scale step sizes per direction) and momentum can escape saddle points: momentum charges through flat regions with inertia, while adaptive methods escape degenerate directions with zero gradient via "per-direction step sizing." Another interesting finding: the "minima" that deep networks find vary widely in performance, and choices like momentum/batch size systematically affect final generalization — this belongs to the "coupling of optimization and generalization," covered in detail in Anatomy of Deep Learning Architectures.

IX. A Brief Introduction to Second-Order Methods ​

First-order methods use only gradients; second-order methods (Newton's method, quasi-Newton, K-FAC, etc.) leverage curvature information (the Hessian matrix):

θ ← θ − H⁻¹·∇L

Theoretically faster convergence and less sensitive to ill-conditioned curvature. But the Hessian is of the square of the parameter count — a billion-parameter model's Hessian has 10¹⁸ elements, which can't be stored or computed. In practice:

  • Newton's method is only feasible for small problems or fine-tuning;
  • Approximate methods like K-FAC have applications in specific scenarios (e.g., reinforcement learning, large-scale distributed training);
  • The practical strategy is "first-order methods + good scheduling/normalization," delegating the curvature problem to Initialization and Normalization for mitigation.

X. Engineering Defaults Table ​

ScenarioDefault RecipeNotes
General small-to-medium CV/NLPAdamW, lr≈3e-4~1e-3, warmup + cosineMost stable and hassle-free
Image classification for SOTASGD + momentum=0.9 + lr=0.1×batch/256 + step/cosineWorks well with BN; convergence metrics are often better
Large model pretraining (>1B)AdamW + warmup + cosine + gradient clipping + BF16 mixed precisionUse Adam betas=(0.9, 0.95)
GANAdam, separate lr for generator/discriminator (e.g., 1e-4/4e-4), gradient penalty optionalTraining is unstable; see VAEs and GANs
Reinforcement LearningAdam + gradient clipping + low lr (1e-4~3e-4)RL is extremely sensitive to hyperparameters; see Deep Reinforcement Learning

The full tuning flow and the decision tree for "which signal to look at when changing which hyperparameter" are covered in Training Recipes and Hyperparameter Tuning.

XI. Tradeoffs ​

Tradeoffs

SGD's generalization vs. Adam's stability: empirically, SGD + momentum often achieves better final generalization (especially in CV), but requires careful scheduling; Adam is nearly plug-and-play and robust, suitable for rapid iteration. A compromise: use Adam first to get the architecture working, then fine-tune with SGD, or just go with AdamW.

Convergence speed vs. generalization quality: large learning rate + large batch converges faster but generalizes slightly worse; small learning rate + small batch generalizes better but is slower. This is the classic "fast vs. good" tradeoff.

The convenience of adaptive step sizing vs. the loss of implicit regularization: Adam independently scales each direction, sacrificing some of SGD's implicit regularization — this is one of the mechanistic explanations for why "Adam generalizes slightly worse."

GPU memory vs. optimization effectiveness: large batch size reduces memory pressure per step but needs more epochs to converge; gradient accumulation can simulate large batches without blowing up GPU memory, but sacrifices some stochasticity.

Optimizers and scheduling are the core of "training recipes," but always remember: optimization can only reach low points that already exist on the loss surface — the loss function and the data determine what that surface looks like. Look at Loss Functions and Output Layers and Data and Data Engineering first, then talk about tuning. For a rapid troubleshooting guide to training failure, see Debugging and Diagnostics.

Further Reading ​

References ​