Skip to content

Backpropagation and Automatic Differentiation

Quick overview Backpropagation is the foundation of deep learning training: the chain rule + gradient backpropagation. This article starts from "why we need gradients," manually derives a small network step by step, then transitions to the two modes of automatic differentiation, and finally covers the correct usage of PyTorch autograd, common pitfalls, and the root cause of vanishing/exploding gradients.

Backpropagation and Automatic Differentiation ​

One-line definition: Backpropagation is the algorithm that efficiently computes the gradient of the loss with respect to every parameter, using the chain rule — all training of deep networks is essentially "backpropagate to compute gradients, then update parameters by descending along them" (see Optimization and Gradient Descent). Automatic differentiation is the systematic, mechanized engineering implementation of this algorithm — the backward() call in PyTorch/TensorFlow runs on it.

I. Why Backpropagation Is Needed ​

Training a neural network requires two things: forward propagation to compute the loss, and backpropagation to compute gradients. We covered forward propagation in Neural Networks Fundamentals: z = Wx + b → activation → … → loss L.

The question: how do you know which direction to adjust each parameter? The answer is to compute ∂L/∂W — the partial derivative of the loss with respect to each weight, i.e., the gradient. Once you have the gradient, parameter updates are straightforward:

W ← W − η · ∂L/∂W

where η is the learning rate. This is gradient descent; details in Optimization and Gradient Descent.

So how do you compute gradients? The dumbest approach is numerical differentiation for each parameter:

∂L/∂Wᵢ ≈ (L(Wᵢ+ε) − L(Wᵢ−ε)) / 2ε

A model with 100 million parameters would need 200 million forward passes — completely impractical. Backpropagation leverages the chain rule to get all gradients in one forward + one backward pass — the computational cost is equivalent to just two forward passes. This is why it dominates deep learning.

Why Backpropagation Matters

Backpropagation was systematized by Rumelhart, Hinton, and Williams in 1986, triggering the deep learning boom. But its core idea (reverse-order chain differentiation) appeared as early as 1970 in Seppo Linnainmaa's paper on automatic differentiation. It reduced the cost of "gradient computation" from O(parameters × forward cost) to O(forward cost), making it possible to "train models with billions of parameters."

II. The Chain Rule: The Mathematical Core ​

The chain rule is a "one-liner" theorem in calculus: the derivative of a composite function equals the product of layer-by-layer derivatives.

If L = f(g(h(x))), then:

∂L/∂x = f′(g(h(x))) · g′(h(x)) · h′(x)

Note the direction: the outermost function is computed first (so f comes first), but when differentiating, we multiply from the outermost L backward. This is the origin of "back" in "backpropagation." For the math fundamentals, see Math Primer.

Applied to a network: the loss L is a function of the output of the last layer, which is a function of the output of the second-to-last layer, and so on. Thus:

∂L/∂W⁽ˡ⁾ = ∂L/∂a⁽ᴸ⁾ · ∂a⁽ᴸ⁾/∂a⁽ᴸ⁻¹⁾ · … · ∂a⁽ˡ⁺¹⁾/∂a⁽ˡ⁾ · ∂a⁽ˡ⁾/∂W⁽ˡ⁾

In a more general form, each intermediate variable z stores an "upstream gradient" ∂L/∂z (also called the gradient signal), and backpropagation propagates this upstream gradient layer by layer through the computation graph.

III. Computation Graphs: Forward and Backward Graphs ​

Any network can be drawn as a computation graph: nodes represent data (tensors) and operations, and edges represent dependencies. Forward propagation computes from input to output in topological order; backpropagation computes gradients from output back to input in reverse topological order.

x ──(Linear)──▶ z ──(ReLU)──▶ a ──(Linear)──▶ ŷ ──(MSE)──▶ L
                    ▲                            ▲
                  W1,b1                        W2,b2

The forward graph stores every intermediate result (z, a) and the local derivatives of each operation. The backward graph uses the chain rule to string these local derivatives together, producing ∂L/∂W1, ∂L/∂W2.

A key engineering point: the memory footprint of the forward pass must be preserved (because the backward pass needs the values of z and a). This is why GPU memory usage during training is roughly equal to "activation values" rather than just parameters — it also explains how memory-saving techniques like gradient checkpointing work.

IV. Manual Derivation of a Small Network: Walk Through It From Scratch ​

Let's manually derive the simplest scalar network: input x=1, one hidden neuron (weight w1=2, no bias), ReLU activation, output weight w2=3, mean squared error loss, target y=10.

Forward propagation:

z1 = w1·x = 2×1 = 2
a1 = ReLU(z1) = 2
ŷ = w2·a1 = 3×2 = 6
L = ½(ŷ − y)² = ½(6−10)² = 8

Backpropagation (from L backward):

∂L/∂ŷ = ŷ − y = 6 − 10 = −4
∂ŷ/∂w2 = a1 = 2            →  ∂L/∂w2 = (−4)×2 = −8
∂ŷ/∂a1 = w2 = 3            →  ∂L/∂a1 = (−4)×3 = −12
∂a1/∂z1 = 1  (ReLU derivative is 1 in the positive region)
∂z1/∂w1 = x = 1            →  ∂L/∂w1 = (−12)×1×1 = −12

So the two gradients are ∂L/∂w2 = −8 and ∂L/∂w1 = −12. With learning rate η=0.1, we get w2 ← 3+0.8 = 3.8 and w1 ← 2+1.2 = 3.2. The loss goes down — this is the minimal closed loop of how a neural network "learns." Generalizing this table from scalars to matrices (each layer has a batch dimension) is exactly what backward() does in real frameworks.

V. Vanishing and Exploding Gradients: The Fate of Multiplication ​

Every step of backpropagation is "multiplication." Look at the chain rule:

∂L/∂W⁽¹⁾ = ∂L/∂a⁽ᴸ⁾ · (Πₗ ∂a⁽ˡ⁺¹⁾/∂a⁽ˡ⁾) · ∂a⁽¹⁾/∂W⁽¹⁾

If the norm of each layer's derivative is less than 1 (e.g., the Sigmoid derivative maxes out at 0.25), after multiplying across 30 layers the gradient ≈ 0.25³⁰, which directly underflows to 0 — vanishing gradients, where shallow-layer parameters barely receive any update signal. Conversely, if every layer's derivative exceeds 1, multiplication leads to exponential explosion — exploding gradients, where parameters jump to NaN in one step.

Engineering Takeaway

90% of "exploding/vanishing gradient" problems in training can be traced to the chain multiplication in backpropagation. Mitigation strategies include: ReLU-family activations, appropriate initialization (Initialization and Normalization), normalization layers, residual connections, and gradient clipping. For troubleshooting, see Debugging and Diagnostics.

A classic victim of vanishing gradients is early RNNs — multiplication across time steps made them nearly incapable of remembering long-range dependencies. This is one of the deep reasons for the later development of LSTM, GRU gating designs, and ultimately the complete replacement of RNNs by Transformers. For the evolutionary narrative, see RNNs and Sequence Modeling.

VI. Automatic Differentiation: Forward Mode vs. Reverse Mode ​

Automatic differentiation is neither numerical differentiation (approximation) nor symbolic differentiation (expanding the expression) — it is the exact application of the chain rule on a computation graph, just automated. It has two basic modes:

  • Forward mode: Starting from the input, simultaneously compute the "directional derivative" for each intermediate variable. Computing the gradient of n inputs requires n forward passes. Suitable for "few inputs, many outputs" (e.g., the gradient of a scalar function with respect to vector parameters), where Jacobian-vector products are efficient.
  • Reverse mode: This is backpropagation. One forward + one backward pass, yielding the gradient of the loss with respect to all parameters. Suitable for the typical neural network scenario of "many inputs, one output" (the loss is scalar, parameters number in the millions) — this is exactly why backpropagation dominates deep learning.

Both modes have their use cases: reverse mode computes all parameter gradients in one shot but must save intermediate activations, which is memory-expensive; forward mode needs no intermediate saves and is memory-friendly, commonly used when gradients themselves participate in computation (e.g., Hessian-vector products). The selection logic and cost tradeoffs are practical considerations covered in How to Choose Frameworks and Tools.

VII. Correct Usage of PyTorch autograd ​

PyTorch's automatic differentiation is powered by torch.autograd. There are just a few core rules:

python
import torch

x = torch.tensor([1.0], requires_grad=True)   # leaf tensor that needs gradients
w = torch.tensor([2.0], requires_grad=True)
b = torch.tensor([0.5], requires_grad=False)  # don't track gradients

z = x * w + b          # operations are recorded
L = z.square().mean()  # loss (scalar)

L.backward()           # backpropagation, fills .grad
print(w.grad)          # tensor([2.0]) — ∂L/∂w

A few semantics you must master:

  • requires_grad=True: This tensor and all tensors derived from it are tracked into the computation graph. Default parameter tensors need gradients; data tensors do not.
  • detach(): "Detaches" the tensor from the computation graph — returns a new tensor that doesn't track gradients and doesn't share recordings. Typical uses: cutting off gradient flow at an intermediate feature (e.g., freezing a feature extractor, stop-gradient in contrastive learning), or disabling gradients on a specific branch outside of torch.no_grad().
  • torch.no_grad(): A context manager where no graphs are built and no activations stored for all operations within the block. Saves GPU memory and time. Must be used during inference, since gradients are not needed.
  • torch.set_grad_enabled(False): Programmatically toggles gradient computation on/off globally. A common pattern in training scripts for switching phases.

The difference between detach() and no_grad() in one sentence: detach() is "a specific tensor exits the graph," while no_grad() is "the entire code block doesn't build a graph."

VIII. Common Pitfalls ​

Backpropagation + autograd is a hotbed of training accidents. Here are the four most common pitfalls:

  1. In-place operations that destroy the computation graph: w += 1, x.add_(1) modify a tensor's value in place, but autograd needs the old values before the operation during the backward pass, which triggers the error a leaf Variable that requires grad is being used in an in-place operation. Correct approach: use w = w + 1 (create a new tensor) or with torch.no_grad(): w.add_(1) for parameter updates.
  2. Updating parameters inside a no_grad() block: with torch.no_grad(): w -= lr * w.grad is valid (parameter updates don't need gradients anyway), but if you execute forward + backward() inside a no_grad() block, gradients are never computed, and the model never updates — this is the #1 reason for "loss stuck for an entire night."
  3. Forgetting to zero gradients each step: loss.backward() accumulates gradients by default. The correct flow is to call optimizer.zero_grad() (or manually zero before loss.backward()) each step, otherwise accumulated gradients cause training oscillation.
  4. Calling backward on non-scalar tensors: backward() requires the output to be scalar (or you must pass gradients of the same shape as the output). For losses with multiple outputs, call .mean() or .sum() first, then backward.

For more troubleshooting of training-time accidents, see Debugging and Diagnostics and Common Pitfalls and Anti-patterns.

IX. Mixed Precision and Gradient Scaling ​

Modern GPUs (such as NVIDIA A100) execute FP16/BF16 operations far faster than FP32, which led to mixed precision training: parameters and primary gradients are stored in FP32, while forward/backward passes use FP16 for speed. But FP16 has a smaller representable range, so gradients may underflow to 0. The solution is gradient scaling:

loss_scaled = loss × scale          # amplify the loss
loss_scaled.backward()              # backward to get amplified gradients
grad = grad / scale                 # restore after backpropagation

PyTorch's torch.cuda.amp.GradScaler automatically handles the flow of "skip the step if overflow, reduce the scale." BF16, which shares the same exponent bits as FP32 and almost loses no range, has become the default choice for LLM training. This entire suite of techniques is virtually standard in production training; see Training Recipes and Hyperparameter Tuning for deployment details.

X. Tradeoffs ​

Tradeoffs

Reverse mode saves compute but costs memory vs. forward mode saves memory but costs compute: use reverse mode for standard training; when memory is the bottleneck and gradient dimensions are low, forward mode or gradient checkpointing is more cost-effective.

Saving all activations vs. checkpointing/recomputation: saving activations is fastest for the backward pass but memory grows linearly with depth; gradient checkpointing stores only a few intermediate values per layer and recomputes during the backward pass, significantly reducing memory at the cost of roughly doubling time — this is the standard "trade time for memory" operation in large model training.

The convenience of autograd vs. controllability: automatic differentiation lets you write models as freely as mathematical formulas, but the "magical" aspect makes problems harder to debug. Understanding every step of the manual example (Section IV) is what gives you the confidence to troubleshoot any gradient issue later.

Automatic differentiation is the foundation of the deep learning edifice: it determines how large a model, how deep a network, and how long a sequence you can train. Master this section, combined with the forward propagation in Neural Networks Fundamentals, and you have all the core pieces of the training loop. The rest is "what to do with gradients," which is covered in the Optimization and Gradient Descent section.

Further Reading ​

References ​