Skip to content

Math Primer

Quick overview A quick-reference math guide for deep learning: linear algebra (vectors/matrices/norms/eigenvalues/SVD), calculus (derivatives/partial derivatives/chain rule/Jacobian/Taylor series), probability & statistics (distributions/expectation/Bayes/MLE/entropy & KL divergence), and optimization (convexity & gradient descent). Each formula includes an explanation of why it matters in deep learning.

Math Primer ​

In one sentence: a quick-reference table organized around "what math DL needs" — four chapters on linear algebra, calculus, probability & statistics, and optimization, each formula paired with "why we need it in deep learning" — for quick lookups during interview follow-ups, paper reading, and hyperparameter tuning.

This is not a systematic textbook. It's a "look-up map": when you encounter an unfamiliar formula, first check the DL context explanation on this page, then confirm the term in the Glossary, and finally see how it's applied on the concept pages.

1. Linear Algebra ​

Vector & Matrix Basics ​

  • Vector: An ordered set of numbers $\mathbf{x} \in \mathbb{R}^n$. In DL, features, embeddings, and single-layer outputs are all vectors.
  • Matrix: $\mathbf{W} \in \mathbb{R}^{m \times n}$, representing a linear transformation. Each layer in a neural network is "matrix multiplication + bias + nonlinearity."
  • Matrix multiplication: $(AB){ij} = \sum_k AB_{kj}$. Forward propagation $h = Wx$ is just $h_i = \sum_j W_{ij}x_j$ — batch processing is just matrix multiplication across many samples at once.
  • Transpose / inverse / identity: $W^T$, $W^{-1}$ (requires invertibility), $I$. Transposes show up frequently during gradient backpropagation (the chain rule's transpose operation is what sends gradients "backwards").
  • Matrix multiplication is associative but not commutative: With $ABC$, you can choose the multiplication order — the order in which you arrange multiplications in the chain rule determines computational cost (e.g., $(AB)x$ vs $A(Bx)$).

Norms & Inner Products ​

  • Inner product: $\langle \mathbf{x}, \mathbf{y}\rangle = \sum_i x_i y_i = \mathbf{x}^T\mathbf{y}$, measuring "how aligned two vectors are." Attention scores are essentially inner products: $q^Tk$.
  • L2 norm: $|\mathbf{x}|_2 = \sqrt{\sum_i x_i^2}$, the Euclidean length. Weight decay (L2 regularization) adds $\lambda|\mathbf{W}|_2^2$ to the loss. See Overfitting & Regularization.
  • L1 norm: $|\mathbf{x}|_1 = \sum_i |x_i|$, which induces sparse solutions. Lasso regularization uses L1.
  • Frobenius norm: $|\mathbf{W}|F = \sqrt{\sum W_{ij}^2}$, the matrix equivalent of the "L2 length."

Eigenvalues & Matrix Decompositions ​

  • Eigenvalues & eigenvectors: $Av = \lambda v$. The "direction of effect" of matrix multiplication is described by eigenvectors. The spectral radius (absolute value of the largest eigenvalue) determines whether linear iterations are stable — this is the mathematical key to understanding gradient explosion / vanishing. See Initialization & Normalization.
  • Eigenvalue decomposition: $A = Q\Lambda Q^{-1}$ (symmetric matrices can be orthogonally diagonalized).
  • Singular Value Decomposition (SVD): $A = U\Sigma V^T$, applicable to any matrix. Use cases: PCA, dimensionality reduction, low-rank matrix approximation, and the "low-rank assumption" behind LoRA (weight increments are approximately low-rank).
  • Positive definite matrix: $\mathbf{x}^TA\mathbf{x} > 0$ (for nonzero x). The bowl-shaped terrain of convex quadratic functions is characterized by positive definite matrices.

2. Calculus ​

Derivatives & Partial Derivatives ​

  • Derivative: $f'(x) = \lim_{h\to 0}\frac{f(x+h)-f(x)}{h}$, the local rate of change of a function at a point.
  • Partial derivative: The derivative of a multivariable function w.r.t. one variable, treating others as constant. The gradient is the vector of all partial derivatives: $\nabla f = (\partial f/\partial x_1, \ldots, \partial f/\partial x_n)$.
  • Common derivatives: $(x^n)' = nx^{n-1}$, $(\log x)' = 1/x$, $(e^x)' = e^x$, $(\sigma(x))' = \sigma(x)(1-\sigma(x))$ (writing the sigmoid derivative as a function of itself is very useful in practice).
  • Softmax derivative: $\partial \text{softmax}i / \partial z_j = p_i(\delta - p_j)$ — this is the origin of the "prediction minus truth" gradient formula. See Loss Functions & Output Layers.

Chain Rule & Automatic Differentiation ​

  • Chain rule: $\frac{dy}{dx} = \frac{dy}{du}\cdot\frac{du}{dx}$. Backpropagation is the engineering implementation of the chain rule: the gradient of a composite function $L = f(g(h(x)))$ equals the product of derivatives of each layer. See Backpropagation & Automatic Differentiation.
  • Jacobian matrix: $J_{ij} = \partial y_i / \partial x_j$, the full derivative of a vector-valued function. The forward/backward transform of each layer can be written as a Jacobian. Residual connections introduce an identity matrix term into the Jacobian, mitigating vanishing gradients.
  • Taylor series: $f(x) \approx f(x_0) + f'(x_0)(x-x_0)$ — gradient descent is repeatedly applying first-order Taylor approximation; the second-order term (Hessian) corresponds to Newton's method. The concept of "loss landscape" and loss terrain visualization derives from this.

Calculus in Gradients & Optimization ​

  • The gradient direction is the direction of steepest function increase, so parameters update along the negative gradient: $\theta \leftarrow \theta - \eta\nabla_\theta L$. See Optimization & Gradient Descent.
  • The learning rate $\eta$ is the step size of the Taylor expansion: too large, and you step out of the reliable region (oscillation/divergence); too small, and convergence is slow.

3. Probability & Statistics ​

Random Variables & Distributions ​

  • Random variable / probability distribution: $X \sim p(x)$, where $p$ describes the likelihood of taking values. A model's output $p(y|x)$ approximates the true conditional distribution.
  • Expectation & variance: $E[X] = \sum_x x p(x)$; $\text{Var}(X) = E[(X-E[X])^2]$. The optimization objective is essentially minimizing expected loss; variance corresponds to the stochastic source of "generalization gap."
  • Common distributions: Bernoulli (binary classification), categorical (multi-class), Gaussian $N(\mu,\sigma^2)$ (regression error, initialization, and noise-adding in diffusion models).
  • IID (Independent and Identically Distributed): The training set samples are assumed IID — breaking this assumption is data drift. See Evaluation in Practice.

Conditional Probability & Bayes ​

  • Conditional probability: $P(A|B) = \frac{P(A\cap B)}{P(B)}$. A classifier's output $P(y|x)$ is a conditional probability.
  • Bayes' theorem: $P(y|x) = \frac{P(x|y)P(y)}{P(x)}$. Updates "prior $P(y)$ + likelihood $P(x|y)$" into "posterior $P(y|x)$." The entire derivation of VAEs is built on "maximizing data likelihood + variational posterior approximation." See VAE & GAN.
  • Maximum Likelihood Estimation (MLE): $\hat\theta = \arg\max_\theta \prod_i p_\theta(x_i)$, equivalent to minimizing negative log-likelihood. Cross-entropy loss is negative log-likelihood — hence cross-entropy is used for classification, which is fundamentally "maximizing the probability of the correct class."

Information Theory ​

  • Entropy: $H(p) = -\sum_x p(x)\log p(x)$, a measure of a distribution's uncertainty.
  • Cross-entropy: $H(p,q) = -\sum_x p(x)\log q(x) = H(p) + D_{KL}(p|q)$. Classification loss = cross-entropy between the true distribution and the predicted distribution.
  • KL divergence: $D_{KL}(p|q) = \sum_x p(x)\log\frac{p(x)}{q(x)}$, measuring distribution difference. It is asymmetric ($D_{KL}(p|q) \neq D_{KL}(q|p)$) and always ≥ 0. KL divergence is ubiquitous in VAEs (variational lower bound), GANs, distillation, and RLHF.

Why does cross-entropy pair with softmax?

Softmax normalizes logits into a probability distribution $q$; the cross-entropy loss derivative w.r.t. logits yields the clean $(p - q)$ form (true minus predicted), with a large and stable gradient. MSE + softmax, on the other hand, has gradients that approach zero in the saturated region, making training extremely slow. See the loss functions and output layers page for details.

4. Optimization ​

  • Convex sets & convex functions: $f(\lambda x + (1-\lambda)y) \le \lambda f(x) + (1-\lambda)f(y)$. Convex functions have no local minima issues (local = global). Neural network losses are non-convex, which is why initialization, LR scheduling, and momentum exist as "engineering weapons for the non-convex era." See Optimization & Gradient Descent.
  • Gradient descent: $\theta_{t+1} = \theta_t - \eta \nabla L(\theta_t)$. Three variants: full-batch (GD), mini-batch (the practice standard), and stochastic (SGD).
  • Momentum: $\theta_{t+1} = \theta_t - \eta v_t$, where $v_t = \beta v_{t-1} + (1-\beta)\nabla L$ — exponential moving average of gradients, suppressing oscillation and accelerating progress along flat directions.
  • Adam: Adapts the learning rate for each parameter (gradient first/second moment), robust to hyperparameters. AdamW corrects the weight decay implementation and is standard for LLM fine-tuning.
  • Saddle points vs local minima: Saddle points (some directions go down, others go up) are far more common than local minima in high-dimensional loss space — this is also why stochasticity and momentum are effective.
  • Learning rate scheduling: warmup, cosine, step decay. Warmup is nearly essential for Transformer pre-training: in early training, parameters are random and gradient noise is high, so start with small steps to stabilize.

5. Quick Reference Table: Essential DL Formulas ​

ConceptFormulaDL Context
Linear layer$h = Wx + b$Core computation of every layer
Attention$\text{softmax}(QK^T/\sqrt{d_k})V$Transformer core. See Transformer Architecture
Cross-entropy loss$L = -\sum_i y_i \log \hat{y}_i$Standard loss for classification tasks
Cross-entropy gradient$\partial L/\partial z = \hat{y} - y$Why training is stable
L2 regularization$L = L_0 + \lambda|\theta|_2^2$Weight decay
Gradient update$\theta \leftarrow \theta - \eta\nabla L$Starting point for all optimization
Bayes' theorem$P(yx) \propto P(x
KL divergence$\sum p\log(p/q)$Generative models and alignment

Further Reading ​

References ​