Theme
Backpropagation and Automatic Differentiation
One-line definition: Backpropagation is the algorithm that efficiently computes the gradient of the loss with respect to every parameter, using the chain rule — all training of deep networks is essentially "backpropagate to compute gradients, then update parameters by descending along them" (see Optimization and Gradient Descent). Automatic differentiation is the systematic, mechanized engineering implementation of this algorithm — the backward() call in PyTorch/TensorFlow runs on it.
I. Why Backpropagation Is Needed
Training a neural network requires two things: forward propagation to compute the loss, and backpropagation to compute gradients. We covered forward propagation in Neural Networks Fundamentals: z = Wx + b → activation → … → loss L.
The question: how do you know which direction to adjust each parameter? The answer is to compute ∂L/∂W — the partial derivative of the loss with respect to each weight, i.e., the gradient. Once you have the gradient, parameter updates are straightforward:
W ← W − η · ∂L/∂Wwhere η is the learning rate. This is gradient descent; details in Optimization and Gradient Descent.
So how do you compute gradients? The dumbest approach is numerical differentiation for each parameter:
∂L/∂Wᵢ ≈ (L(Wᵢ+ε) − L(Wᵢ−ε)) / 2εA model with 100 million parameters would need 200 million forward passes — completely impractical. Backpropagation leverages the chain rule to get all gradients in one forward + one backward pass — the computational cost is equivalent to just two forward passes. This is why it dominates deep learning.
Why Backpropagation Matters
Backpropagation was systematized by Rumelhart, Hinton, and Williams in 1986, triggering the deep learning boom. But its core idea (reverse-order chain differentiation) appeared as early as 1970 in Seppo Linnainmaa's paper on automatic differentiation. It reduced the cost of "gradient computation" from O(parameters × forward cost) to O(forward cost), making it possible to "train models with billions of parameters."
II. The Chain Rule: The Mathematical Core
The chain rule is a "one-liner" theorem in calculus: the derivative of a composite function equals the product of layer-by-layer derivatives.
If L = f(g(h(x))), then:
∂L/∂x = f′(g(h(x))) · g′(h(x)) · h′(x)Note the direction: the outermost function is computed first (so f comes first), but when differentiating, we multiply from the outermost L backward. This is the origin of "back" in "backpropagation." For the math fundamentals, see Math Primer.
Applied to a network: the loss L is a function of the output of the last layer, which is a function of the output of the second-to-last layer, and so on. Thus:
∂L/∂W⁽ˡ⁾ = ∂L/∂a⁽ᴸ⁾ · ∂a⁽ᴸ⁾/∂a⁽ᴸ⁻¹⁾ · … · ∂a⁽ˡ⁺¹⁾/∂a⁽ˡ⁾ · ∂a⁽ˡ⁾/∂W⁽ˡ⁾In a more general form, each intermediate variable z stores an "upstream gradient" ∂L/∂z (also called the gradient signal), and backpropagation propagates this upstream gradient layer by layer through the computation graph.
III. Computation Graphs: Forward and Backward Graphs
Any network can be drawn as a computation graph: nodes represent data (tensors) and operations, and edges represent dependencies. Forward propagation computes from input to output in topological order; backpropagation computes gradients from output back to input in reverse topological order.
x ──(Linear)──▶ z ──(ReLU)──▶ a ──(Linear)──▶ ŷ ──(MSE)──▶ L
▲ ▲
W1,b1 W2,b2The forward graph stores every intermediate result (z, a) and the local derivatives of each operation. The backward graph uses the chain rule to string these local derivatives together, producing ∂L/∂W1, ∂L/∂W2.
A key engineering point: the memory footprint of the forward pass must be preserved (because the backward pass needs the values of z and a). This is why GPU memory usage during training is roughly equal to "activation values" rather than just parameters — it also explains how memory-saving techniques like gradient checkpointing work.
IV. Manual Derivation of a Small Network: Walk Through It From Scratch
Let's manually derive the simplest scalar network: input x=1, one hidden neuron (weight w1=2, no bias), ReLU activation, output weight w2=3, mean squared error loss, target y=10.
Forward propagation:
z1 = w1·x = 2×1 = 2
a1 = ReLU(z1) = 2
ŷ = w2·a1 = 3×2 = 6
L = ½(ŷ − y)² = ½(6−10)² = 8Backpropagation (from L backward):
∂L/∂ŷ = ŷ − y = 6 − 10 = −4
∂ŷ/∂w2 = a1 = 2 → ∂L/∂w2 = (−4)×2 = −8
∂ŷ/∂a1 = w2 = 3 → ∂L/∂a1 = (−4)×3 = −12
∂a1/∂z1 = 1 (ReLU derivative is 1 in the positive region)
∂z1/∂w1 = x = 1 → ∂L/∂w1 = (−12)×1×1 = −12So the two gradients are ∂L/∂w2 = −8 and ∂L/∂w1 = −12. With learning rate η=0.1, we get w2 ← 3+0.8 = 3.8 and w1 ← 2+1.2 = 3.2. The loss goes down — this is the minimal closed loop of how a neural network "learns." Generalizing this table from scalars to matrices (each layer has a batch dimension) is exactly what backward() does in real frameworks.
V. Vanishing and Exploding Gradients: The Fate of Multiplication
Every step of backpropagation is "multiplication." Look at the chain rule:
∂L/∂W⁽¹⁾ = ∂L/∂a⁽ᴸ⁾ · (Πₗ ∂a⁽ˡ⁺¹⁾/∂a⁽ˡ⁾) · ∂a⁽¹⁾/∂W⁽¹⁾If the norm of each layer's derivative is less than 1 (e.g., the Sigmoid derivative maxes out at 0.25), after multiplying across 30 layers the gradient ≈ 0.25³⁰, which directly underflows to 0 — vanishing gradients, where shallow-layer parameters barely receive any update signal. Conversely, if every layer's derivative exceeds 1, multiplication leads to exponential explosion — exploding gradients, where parameters jump to NaN in one step.
Engineering Takeaway
90% of "exploding/vanishing gradient" problems in training can be traced to the chain multiplication in backpropagation. Mitigation strategies include: ReLU-family activations, appropriate initialization (Initialization and Normalization), normalization layers, residual connections, and gradient clipping. For troubleshooting, see Debugging and Diagnostics.
A classic victim of vanishing gradients is early RNNs — multiplication across time steps made them nearly incapable of remembering long-range dependencies. This is one of the deep reasons for the later development of LSTM, GRU gating designs, and ultimately the complete replacement of RNNs by Transformers. For the evolutionary narrative, see RNNs and Sequence Modeling.
VI. Automatic Differentiation: Forward Mode vs. Reverse Mode
Automatic differentiation is neither numerical differentiation (approximation) nor symbolic differentiation (expanding the expression) — it is the exact application of the chain rule on a computation graph, just automated. It has two basic modes:
- Forward mode: Starting from the input, simultaneously compute the "directional derivative" for each intermediate variable. Computing the gradient of
ninputs requiresnforward passes. Suitable for "few inputs, many outputs" (e.g., the gradient of a scalar function with respect to vector parameters), where Jacobian-vector products are efficient. - Reverse mode: This is backpropagation. One forward + one backward pass, yielding the gradient of the loss with respect to all parameters. Suitable for the typical neural network scenario of "many inputs, one output" (the loss is scalar, parameters number in the millions) — this is exactly why backpropagation dominates deep learning.
Both modes have their use cases: reverse mode computes all parameter gradients in one shot but must save intermediate activations, which is memory-expensive; forward mode needs no intermediate saves and is memory-friendly, commonly used when gradients themselves participate in computation (e.g., Hessian-vector products). The selection logic and cost tradeoffs are practical considerations covered in How to Choose Frameworks and Tools.
VII. Correct Usage of PyTorch autograd
PyTorch's automatic differentiation is powered by torch.autograd. There are just a few core rules:
python
import torch
x = torch.tensor([1.0], requires_grad=True) # leaf tensor that needs gradients
w = torch.tensor([2.0], requires_grad=True)
b = torch.tensor([0.5], requires_grad=False) # don't track gradients
z = x * w + b # operations are recorded
L = z.square().mean() # loss (scalar)
L.backward() # backpropagation, fills .grad
print(w.grad) # tensor([2.0]) — ∂L/∂wA few semantics you must master:
requires_grad=True: This tensor and all tensors derived from it are tracked into the computation graph. Default parameter tensors need gradients; data tensors do not.detach(): "Detaches" the tensor from the computation graph — returns a new tensor that doesn't track gradients and doesn't share recordings. Typical uses: cutting off gradient flow at an intermediate feature (e.g., freezing a feature extractor, stop-gradient in contrastive learning), or disabling gradients on a specific branch outside oftorch.no_grad().torch.no_grad(): A context manager where no graphs are built and no activations stored for all operations within the block. Saves GPU memory and time. Must be used during inference, since gradients are not needed.torch.set_grad_enabled(False): Programmatically toggles gradient computation on/off globally. A common pattern in training scripts for switching phases.
The difference between detach() and no_grad() in one sentence: detach() is "a specific tensor exits the graph," while no_grad() is "the entire code block doesn't build a graph."
VIII. Common Pitfalls
Backpropagation + autograd is a hotbed of training accidents. Here are the four most common pitfalls:
- In-place operations that destroy the computation graph:
w += 1,x.add_(1)modify a tensor's value in place, but autograd needs the old values before the operation during the backward pass, which triggers the errora leaf Variable that requires grad is being used in an in-place operation. Correct approach: usew = w + 1(create a new tensor) orwith torch.no_grad(): w.add_(1)for parameter updates. - Updating parameters inside a
no_grad()block:with torch.no_grad(): w -= lr * w.gradis valid (parameter updates don't need gradients anyway), but if you execute forward +backward()inside ano_grad()block, gradients are never computed, and the model never updates — this is the #1 reason for "loss stuck for an entire night." - Forgetting to zero gradients each step:
loss.backward()accumulates gradients by default. The correct flow is to calloptimizer.zero_grad()(or manually zero beforeloss.backward()) each step, otherwise accumulated gradients cause training oscillation. - Calling backward on non-scalar tensors:
backward()requires the output to be scalar (or you must pass gradients of the same shape as the output). For losses with multiple outputs, call.mean()or.sum()first, then backward.
For more troubleshooting of training-time accidents, see Debugging and Diagnostics and Common Pitfalls and Anti-patterns.
IX. Mixed Precision and Gradient Scaling
Modern GPUs (such as NVIDIA A100) execute FP16/BF16 operations far faster than FP32, which led to mixed precision training: parameters and primary gradients are stored in FP32, while forward/backward passes use FP16 for speed. But FP16 has a smaller representable range, so gradients may underflow to 0. The solution is gradient scaling:
loss_scaled = loss × scale # amplify the loss
loss_scaled.backward() # backward to get amplified gradients
grad = grad / scale # restore after backpropagationPyTorch's torch.cuda.amp.GradScaler automatically handles the flow of "skip the step if overflow, reduce the scale." BF16, which shares the same exponent bits as FP32 and almost loses no range, has become the default choice for LLM training. This entire suite of techniques is virtually standard in production training; see Training Recipes and Hyperparameter Tuning for deployment details.
X. Tradeoffs
Tradeoffs
Reverse mode saves compute but costs memory vs. forward mode saves memory but costs compute: use reverse mode for standard training; when memory is the bottleneck and gradient dimensions are low, forward mode or gradient checkpointing is more cost-effective.
Saving all activations vs. checkpointing/recomputation: saving activations is fastest for the backward pass but memory grows linearly with depth; gradient checkpointing stores only a few intermediate values per layer and recomputes during the backward pass, significantly reducing memory at the cost of roughly doubling time — this is the standard "trade time for memory" operation in large model training.
The convenience of autograd vs. controllability: automatic differentiation lets you write models as freely as mathematical formulas, but the "magical" aspect makes problems harder to debug. Understanding every step of the manual example (Section IV) is what gives you the confidence to troubleshoot any gradient issue later.
Automatic differentiation is the foundation of the deep learning edifice: it determines how large a model, how deep a network, and how long a sequence you can train. Master this section, combined with the forward propagation in Neural Networks Fundamentals, and you have all the core pieces of the training loop. The rest is "what to do with gradients," which is covered in the Optimization and Gradient Descent section.
Further Reading
- Neural Networks Fundamentals — forward propagation and network structure
- Anatomy of Deep Learning Architectures — a panoramic view of the four pillars: data, models, loss, optimization
- Initialization and Normalization — the other half of mitigating vanishing/exploding gradients
- Loss Functions and Output Layers — the quality source of gradient signals
- Deep Learning Evaluation and Experiments — gradient diagnostics in experiments and hyperparameter tuning
- Math Primer — mathematical foundations: chain rule, Jacobians, etc.
References
- Rumelhart, Hinton, Williams. Learning representations by back-propagating errors (Nature, 1986)
- Linnainmaa. The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors (1970)
- Baydin, Pearlmutter, Radul, Siskind. Automatic Differentiation in Machine Learning: a Survey (2018)
- Paszke et al. Automatic differentiation in PyTorch (2017)
- PyTorch autograd Documentation