Theme
Math Primer
Machine learning depends on math, yet countless beginners treat it as a wall they must "climb before proceeding": finish a full course in advanced calculus, linear algebra, and probability theory, then start ML. The result is often the same — three months later, half the math is forgotten and ML hasn't even begun.
This article tears that wall down. Its approach aligns with our site's consistent philosophy — math is not a barrier, it's a tool: every piece of math you'll actually use maps to a concrete machine learning scenario. This article reorganizes the math that ML truly requires by "use case", so you can see clearly what problem each concept solves, where it appears in which models, and where to learn it.
How to use this article
This isn't a textbook — it's a "map + dictionary". Read through it once for the big picture, then return to the relevant section when you encounter an unfamiliar formula. It complements the Glossary: the glossary covers "what" something is, while this article explains "why".
I. Why Machine Learning Needs Math: Not a Barrier, But a Tool
1.1 What Machine Learning Actually Does
Compressing machine learning into one sentence: from a vast space of functions, find one that performs as well as possible on your data. Unpacked, this breaks into three steps:
- Data are points: each sample is a point (vector) in high-dimensional space; features are its coordinates.
- Models are functions: a model is essentially a parameterized function $y = f(x; \theta)$, where parameters $\theta$ determine the function's shape.
- Learning is searching: continuously adjust $\theta$ so that $f$ minimizes error on training data while still performing well on unseen data (generalization).
These three steps map neatly onto three pillars of foundational math:
| Step | Math Tool | Question It Answers | Typical Scenario |
|---|---|---|---|
| Describe data | Linear algebra | What does data look like, how far apart are points? | Feature vectors, similarity, dimensionality reduction |
| Tune parameters | Calculus | Which direction should parameters change? | Gradient descent, backpropagation |
| Define "good" vs "bad" | Probability & stats + Information theory | Why does the loss function look like this? | Cross-entropy, maximum likelihood, regularization |
Add a fourth pillar — optimization — as the engine: gradient descent is its gateway. We'll unpack each one below.
1.2 Why "Finish Math Before ML" Is the Wrong Approach
The mathematical system is vast enough to fill a lifetime, while ML actually uses only a highly concentrated sliver of it. Even more importantly, modern deep learning frameworks (PyTorch, TensorFlow, JAX) all include automatic differentiation — you hardly ever need to manually derive gradients.
Common anti-pattern
"Study a full calculus course before ML" is the least efficient path: without a concrete problem to anchor it, math knowledge is neither memorable nor applicable. The correct approach is the opposite: let ML concepts come to you first (overfitting, loss functions, gradient descent…), and when there's math lurking behind them, return to the relevant section of this article. See Learning Paths and What Is Machine Learning for recommended entry sequences.
1.3 How Much You Need to "Know"
One sentence: be able to compute, understand the intuition, and read formulas — but proofs aren't required.
- Be able to compute: given a small matrix, perform matrix multiplication by hand; given a simple loss function, derive one step of the gradient by hand.
- Understand the intuition: know what "gradient points in the steepest direction" and "larger eigenvalue = more variance in that direction" actually mean.
- Be able to read formulas: see $\sum$, $\nabla$, $\mathbb{E}$ in a paper without panicking, and understand what they're computing.
- Proofs aren't required: you don't need to derive theorems from axioms. A machine learning engineer's job is to turn concepts into code, not publish math papers.
The five chapters below are ordered by "frequency of use", each ending with a set of real, verifiable self-study resources.
II. Linear Algebra: The Universal Language of Data and Models
If machine learning has a native language, it's linear algebra. Data are vectors, model parameters are matrices, and a neural network's forward pass is a chain of matrix multiplications.
2.1 Vectors: Translating Data into Geometry
A single sample is a vector: with $d$ features, it lives in $d$-dimensional space.
$$ x = (x_1, x_2, \dots, x_d) $$
Take house price prediction as an example: a house's features might be $(area=85, bedrooms=3, age=12)$, which is a point in 3D space. A vector has two identities:
- Algebraic identity: an ordered list of numbers.
- Geometric identity: a directed line segment from the origin to that point, with both direction and magnitude.
These two identities are the source of all intuition: "similarity" geometrically means "directions are close", and closeness of direction is measured by the dot product (see §2.4).
2.2 Matrices and Matrix Multiplication
A matrix is a rectangular array of numbers. In ML it has two common readings, both widely used:
- Data table: one row per sample, one column per feature. Any tabular data (DataFrame) you encounter is a matrix.
- Linear transformation: mapping one vector to another, producing scaling, rotation, shearing (or a combination).
The key to matrix multiplication is that "each element of the result matrix is a dot product". If $C = AB$, then
$$ C_{ij} = \sum_{k} A_{ik} B_{kj} $$
That is, the element at row $i$ and column $j$ of $C$ equals the element-wise product of row $i$ of $A$ and column $j$ of $B$, summed. This is precisely weighted sum — and the weighted sum is the sole computation unit of neural networks.
2.3 Why Matrix Multiplication Is Neural Networks
Consider one layer of a fully connected neural network:
$$ h = \sigma(W x + b) $$
- $x$ is the input vector (output from the previous layer).
- $W$ is the weight matrix; each component of $Wx$ is the "dot product of the input with a row of $W$", i.e., a weighted sum for the neuron corresponding to that row.
- $b$ is the bias; $\sigma$ is the activation function (introducing nonlinearity).
One layer of a network = "one linear transformation $Wx+b$, followed by an element-wise nonlinear function $\sigma$". A multi-layer network = multiple such compositions nested. So "deep learning", from a mathematical perspective, is simply nested composite functions. Stacking matrix multiplications is literally linear algebra's native craft. A more systematic treatment is in Deep Learning Fundamentals.
2.4 Dot Product and Similarity
The dot product of two vectors $a, b$ is defined as
$$ a \cdot b = \sum_i a_i b_i = |a| \cdot |b| \cdot \cos\theta $$
where $\theta$ is the angle between the two vectors. This equation reveals the geometric meaning of the dot product: it measures how aligned the directions of two vectors are — maximized when directions are the same, zero when perpendicular.
This gives us ML's most commonly used similarity measure — cosine similarity:
$$ \cos\theta = \frac{a \cdot b}{|a| \cdot |b|} $$
It only cares about direction, unaffected by vector magnitude (e.g., term frequency). This is the default metric for recommendation systems, vector retrieval, semantic search, and Embedding comparisons: encode users, items, and sentences into vectors, and similarity is the "cosine of the angle between user vector and item vector". It even has its own name for the computation — dot product after normalization.
2.5 Norms: Measuring Vector Magnitude
A norm is a generalization of "vector length". The two most common are:
- L2 norm (Euclidean length): $|x|_2 = \sqrt{x_1^2 + \dots + x_d^2}$, corresponding to "straight-line distance".
- L1 norm: $|x|_1 = |x_1| + \dots + |x_d|$, corresponding to "Manhattan distance" (moving along coordinate axes).
They appear in two critical places:
- Distance measurement: KNN and K-Means use Euclidean distance to find the "nearest samples".
- Regularization: L2 regularization applies a $|w|_2^2$ penalty to the parameter vector, L1 applies $|w|_1$ — the former shrinks all weights overall, the latter forces some weights to precisely zero out (feature selection). See Regularization for details.
2.6 Eigenvalue Decomposition: The Heart of PCA
Eigenvalues and eigenvectors answer a profound question: what "directions stay unchanged" under a linear transformation? If there exists a non-zero vector $v$ and scalar $\lambda$ such that
$$ A v = \lambda v $$
then $v$ is an eigenvector of $A$, and $\lambda$ is the corresponding eigenvalue. Geometrically: applying transformation $A$ to $v$ does not change its direction, only scales its length by $\lambda$. Larger eigenvalues mean that direction was stretched more.
Principal Component Analysis (PCA) is directly tied to eigenvalue decomposition. PCA's goal is: find a set of directions onto which, when you project the data, the variance is maximized (most information preserved). It can be shown that these directions are precisely the eigenvectors of the data covariance matrix, and the projected variances are the corresponding eigenvalues. Thus:
$$ \text{Data} \xrightarrow{\text{compute covariance}} \Sigma \xrightarrow{\text{eigen decompose}} \Sigma v = \lambda v \xrightarrow{\text{keep top } k \text{ eigenvalues}} \text{Principal components} $$
Intuition: sorting eigenvalues is "sorting by information content". Keeping the top $k$ eigenvectors to form a projection matrix reduces $d$-dimensional data to $k$ dimensions while losing the least information. This is the standard approach for dimensionality reduction, denoising, and visualization (reducing to 2/3 dimensions).
2.7 Concept → ML Use Case Quick Reference
| Linear Algebra Concept | ML Use Case | Where It Appears |
|---|---|---|
| Vector | Data representation | Everything: features, Embeddings |
| Matrix multiplication | Weighted sums, feature combinations | Every layer of a neural network |
| Dot product / cosine similarity | Measuring similarity | Retrieval, recommendations, clustering |
| Norms | Measuring magnitude, regularization | Distance-based algorithms, L1/L2 regularization |
| Eigenvalues / eigenvectors | Finding directions of maximum variance | PCA, spectral clustering, PageRank |
| Matrix inversion | Analytical solutions | Closed-form solution for least squares |
| Determinant | Measuring volume / invertibility | Theoretical analysis (rarely used directly) |
| Transpose | Dimension alignment | Any matrix operation |
Self-Study Resources · Linear Algebra
- 3Blue1Brown's The Essence of Linear Algebra (YouTube): must-watch, ~15 short videos that turn vectors, matrix multiplication, and eigenvalues into geometric animations. One video per topic, ~10 minutes each, building all intuition from the ground up.
- Gilbert Strang's MIT 18.06 Linear Algebra: free full course on MIT OpenCourseWare, with companion textbook Introduction to Linear Algebra (Wellesley-Cambridge Press) — the "canonical yet approachable" path for linear algebra.
- Khan Academy Linear Algebra: many step-by-step exercises, great for hands-on reinforcement.
III. Calculus: Finding "Which Direction to Adjust"
Linear algebra describes the world; calculus handles "change". The learning process in ML is constantly tweaking parameters, and calculus provides the direction: "which way to adjust".
3.1 Derivatives: Instantaneous Rate of Change
The derivative of a single-variable function $f(x)$ at $x_0$ is
$$ f'(x_0) = \lim_{h \to 0} \frac{f(x_0 + h) - f(x_0)}{h} $$
It represents the "instantaneous rate of change as the input varies" at $x_0$, geometrically the slope of the tangent line. In ML, the derivative of a loss function $L(w)$ answers: if $w$ increases slightly, does the loss go up or down, and by how much? The sign of the derivative tells us which direction to move $w$ to make $L$ decrease.
3.2 Partial Derivatives and the Gradient
Parameters usually number many ($w_1, \dots, w_d$), and the loss function $L(w_1, \dots, w_d)$ is multivariate. Taking the derivative with respect to a single variable while holding others fixed gives the partial derivative $\frac{\partial L}{\partial w_i}$. Stack all partial derivatives into a vector, and you get the gradient:
$$ \nabla L = \left( \frac{\partial L}{\partial w_1}, \dots, \frac{\partial L}{\partial w_d} \right) $$
The gradient has two core properties, the starting point for all optimization intuition:
- The gradient's direction is where the function "rises most steeply".
- The gradient's magnitude is the rate of that rise.
So to decrease the loss, move parameters in the negative gradient direction — this is gradient descent (unpacked in Chapter 6).
3.3 Chain Rule = Backpropagation
Neural networks are nested functions, and their gradients are computed by one rule — the chain rule:
$$ \frac{dL}{dx} = \frac{dL}{dh} \cdot \frac{dh}{dz} \cdot \frac{dz}{dx} $$
If $y = f(g(h(x)))$, then the derivative with respect to $x$ equals the product of each layer's derivative. Between a neural network's parameters $\theta$ and the loss $L$, there are dozens of layers of functions. The chain rule tells us: multiply each layer's local derivative, and you can trace all the way from the output back to the input.
Backpropagation is the engineering implementation of the chain rule:
Loss L ──→ ∂L/∂h_output ──→ ∂L/∂W_last ──→ …… ──→ ∂L/∂W_first
(propagated backward through the computation graph)It first performs a forward pass to compute each layer's output, then from the output layer, computes gradients layer by layer in reverse, caching intermediate results to avoid repeated computation. You don't need to implement it by hand today — the framework's autograd is its implementation — but understanding that "gradients propagate backward from the loss along the computation graph" is crucial for debugging networks (e.g., vanishing gradients, where the product of derivatives approaches zero). See Deep Learning Fundamentals for details.
3.4 Taylor Expansion: Approximating Complex Functions Locally
Taylor's theorem states: a smooth function near a point can be approximated by a polynomial. Keeping only the first-order term gives linear approximation:
$$ f(x) \approx f(x_0) + f'(x_0)(x - x_0) $$
This looks trivial, but it's the mathematical origin of gradient descent. Doing a first-order expansion of the loss function near $w_t$ and finding "the direction that decreases $L$ fastest" yields precisely the negative gradient $- \nabla L$ — a "take one small step" linearized reasoning. Keeping the second-order term gives
$$ f(x) \approx f(x_0) + f'(x_0)(x - x_0) + \frac{1}{2} f''(x_0)(x - x_0)^2 $$
The second-order term uses second derivatives (the Hessian matrix), corresponding to Newton's method family of optimizers — rarely used in practice (computationally expensive), but worth knowing exists.
3.5 Connection Points to ML
| Calculus Concept | ML Use Case |
|---|---|
| Derivative | Determining which direction to adjust a parameter for the loss to decrease |
| Gradient | Direction of steepest increase of the loss function (negative gradient = direction of decrease) |
| Chain rule | Backpropagation = gradient computation for multi-layer composite functions |
| Partial derivatives | Computing gradients per-parameter (high-dimensional) |
| Taylor expansion | Local linearization basis for gradient descent and learning rate choices |
| Integration (rare) | Probability normalization, expectation computation (theoretical side) |
Self-Study Resources · Calculus
- 3Blue1Brown's The Essence of Calculus: top-tier visualizations for derivative and integral intuition.
- 3Blue1Brown's neural network series: especially episode 3, which turns gradient descent and backpropagation into animations — the best demonstration of "how calculus serves ML".
- Khan Academy Multivariable Calculus: systematic courses on partial derivatives, gradients, and the chain rule.
- MIT OCW 18.02 Multivariable Calculus (taught by Strang/Edwards): complete systematic study.
IV. Probability and Statistics: Modeling Uncertainty
Real-world data comes with noise; model predictions carry uncertainty. Probability theory is the language that describes this uncertainty; statistical inference answers "what can we deduce from the data". ML loss functions, evaluation metrics, and model assumptions all grow from here.
4.1 Axioms of Probability
Two axioms of probability, worth remembering intuitively:
- Any event's probability lies in $[0, 1]$, and the total probability of the sample space is 1.
- Mutually exclusive events add: $P(A \cup B) = P(A) + P(B)$ (when mutually exclusive).
This ensures that "probability" is a consistent measurement system, not arbitrary percentages. In ML, model outputs of "confidence" and classification probabilities all obey these rules.
4.2 Conditional Probability and Bayes' Theorem
Conditional probability $P(A \mid B)$ means "the probability of A given that B has occurred". Its definition is
$$ P(A \mid B) = \frac{P(A \cap B)}{P(B)} $$
From this, Bayes' theorem follows — one of the most important equations in ML:
$$ P(\theta \mid D) = \frac{P(D \mid \theta), P(\theta)}{P(D)} $$
Each component has a standard name and intuition:
- Prior $P(\theta)$: your belief about the parameter/hypothesis before seeing data (e.g., "the prior probability of spam is 20%").
- Likelihood $P(D \mid \theta)$: given the parameter, how likely is the data?
- Posterior $P(\theta \mid D)$: belief updated after seeing data — this is the probabilistic definition of "learning".
- Evidence $P(D)$: the marginal probability of the data, usually serving only as a normalization constant.
One sentence: posterior ∝ likelihood × prior. Learning = updating the prior with data (likelihood). Naive Bayes classifiers, Bayesian networks, and another understanding of regularization (Gaussian prior → L2 regularization) all stand on its shoulders.
4.3 Random Variables and Common Distributions
A random variable maps outcomes of random experiments to numbers. It follows some distribution, describing the probability pattern of its values. Three distributions plus one supporting actor repeatedly appear in ML:
| Distribution | Notation | Values | ML Appearance |
|---|---|---|---|
| Bernoulli | $\mathrm{Bern}(p)$ | 0 or 1 | Binary classification: $P(y=1)=p$ |
| Categorical | $\mathrm{Cat}(p_1,\dots,p_K)$ | $1,\dots,K$ | Multi-class: softmax output |
| Normal (Gaussian) | $\mathcal{N}(\mu, \sigma^2)$ | All reals | Noise assumption, regression error, weight initialization, feature distribution |
| Uniform | $\mathrm{Unif}(a,b)$ | Within interval | Random parameter initialization |
The normal distribution deserves extra attention: its probability density function
$$ p(x) = \frac{1}{\sqrt{2\pi\sigma^2}} \exp\left( -\frac{(x-\mu)^2}{2\sigma^2} \right) $$
is symmetric around mean $\mu$, with $\sigma$ controlling spread. "Errors are normally distributed" is the most common assumption in regression problems, and it is precisely the origin of the MSE loss in §4.5 (not some other loss).
4.4 Expectation and Variance
- Expectation $\mathbb{E}[X] = \sum x, p(x)$ (discrete), is the center of the distribution (weighted average).
- Variance $\mathrm{Var}(X) = \mathbb{E}[(X - \mathbb{E}[X])^2]$, measures the magnitude of fluctuations.
In the evaluation section, they combine into the most famous equation — the bias-variance decomposition. For regression error (MSE):
$$ \mathbb{E}[(y - \hat{f}(x))^2] = \underbrace{\text{Bias}^2}{\text{Bias}^2} + \underbrace{\text{Variance}}{\text{Variance}} + \underbrace{\sigma^2}_{\text{Irreducible noise}} $$
Intuition: error = model's systematic bias + model's oversensitivity to the training set (variance) + the data's inherent noise. Simpler models have higher bias and lower variance; more complex models have the opposite — this is the mathematical root of overfitting and underfitting, and what regularization aims to tune. Full discussion in Model Evaluation and Validation and Regularization.
4.5 Maximum Likelihood Estimation: Where Loss Functions Come From
This is one of the most important derivations in this entire article. Maximum Likelihood Estimation (MLE): choose the parameter $\theta^*$ that makes the "probability of the observed data maximum".
$$ \theta^* = \arg\max_\theta \prod_{i} p(y_i \mid x_i; \theta) $$
Products are hard to compute, so take the logarithm (a monotonic function that doesn't change the optimum), turning products into sums:
$$ \theta^* = \arg\max_\theta \sum_i \log p(y_i \mid x_i; \theta) $$
Now, making two different assumptions about the likelihood yields two different loss functions:
Assumption 1: errors follow a normal distribution $y = f_\theta(x) + \varepsilon,\ \varepsilon \sim \mathcal{N}(0,\sigma^2)$. Substituting into the log-likelihood and simplifying (dropping constants), maximizing the likelihood is equivalent to
$$ \arg\min_\theta \sum_i \left( y_i - f_\theta(x_i) \right)^2 $$
— this is the Mean Squared Error (MSE) loss! Regression loss wasn't pulled from thin air; it's the maximum likelihood under the "errors are normal" assumption.
Assumption 2: binary classification output follows a Bernoulli distribution $y \sim \mathrm{Bern}(p_\theta(x))$. Maximizing the log-likelihood simplifies to an equivalent form of
$$ \arg\min_\theta -\sum_i \left[ y_i \log \hat{p}_i + (1 - y_i)\log(1 - \hat{p}_i) \right] $$
— this is the cross-entropy loss (binary cross-entropy for binary classification). It is also the same expression derived in Chapter 5 from information theory: two paths to the same loss, showing that cross-entropy both "minimizes encoding cost" and "maximizes likelihood".
So when someone asks "why cross-entropy for classification and MSE for regression", the answer is: they are MLE under Bernoulli / Gaussian assumptions. This perspective dissolves countless confusions instantly.
4.6 Central Limit Theorem: Why Noise Is Always Normal
Central Limit Theorem (CLT), intuitive version: the sum of many independent, ident distributed small random variables approaches a normal distribution, regardless of what distribution each small variable itself follows.
$$ \frac{1}{n}\sum_{i=1}^n X_i ;\xrightarrow{n \to \infty}; \mathcal{N}\left(\mu, \frac{\sigma^2}{n}\right) $$
The intuitive value is enormous: measurement errors are superpositions of many small errors → approximately normal; average estimates $\bar{x}$ fluctuate around the true value → we can compute confidence intervals. It also explains why "assuming errors are normal" in §4.5 makes sense — normal is not an arbitrary choice; it's nature's default outcome. In ML, it also supports the intuition behind using bootstrap, hypothesis testing, and other statistical tools.
Self-Study Resources · Probability and Statistics
- Khan Academy Probability and Statistics: comprehensive introduction to axioms, conditional probability, Bayes, and distributions.
- Harvard Stat 110 Introduction to Probability (Joe Blitzstein, YouTube/edX free): the best-reviewed probability course, balancing intuition and rigor.
- MIT OCW 18.05 Introduction to Probability and Statistics: more applied.
- StatQuest (Josh Starmer, YouTube): visual quick-takes on Bayes, p-values, confidence intervals, and other concepts — great for review.
V. Information Theory: The Root of Loss Functions
Information theory studies "how to quantify information", and it gives machine learning two gifts: entropy (quantifying uncertainty) and KL divergence (quantifying the distance between two distributions). The "standard answer" for classification loss functions grows from here.
5.1 Information Content and Entropy
The information content an event carries is inversely proportional to its probability: a certain event (probability 1) carries no information, while a very rare event (probability near 0) carries enormous information.
$$ I(x) = -\log_2 p(x) $$
Averaging over a random variable gives entropy:
$$ H(p) = -\sum_x p(x) \log_2 p(x) $$
Intuition: entropy is "how many bits, on average, are needed to encode samples from this distribution" — i.e., a measure of uncertainty. A fair coin flip ($p=0.5$) has entropy of 1 bit; a certain event has 0 entropy; a uniform distribution has maximum entropy (most uncertain). In ML, when a model predicts "this class is a cat" with $p=0.99$, uncertainty is low — small entropy; with $p=0.33$, it's full of uncertainty — large entropy.
5.2 Cross-Entropy
If the true distribution is $p$, but you use a (possibly incorrect) distribution $q$ to encode it, the average number of bits needed is the cross-entropy:
$$ H(p, q) = -\sum_x p(x) \log_2 q(x) $$
Cross-entropy is always greater than or equal to entropy ($H(p,q) \ge H(p)$); the extra part is the "cost of using the wrong encoding".
5.3 KL Divergence: Distance Between Two Distributions
Taking the difference between cross-entropy and entropy separately gives KL divergence (Kullback–Leibler divergence):
$$ D_{KL}(p | q) = \sum_x p(x) \log_2 \frac{p(x)}{q(x)} = H(p,q) - H(p) $$
- $D_{KL}(p | q) \ge 0$, and $D_{KL} = 0$ if and only if $p = q$;
- It is the "cost of approximating $p$ with $q$", commonly used to compare predicted distributions with true distributions.
- Note it is asymmetric: $D_{KL}(p|q) \ne D_{KL}(q|p)$, so strictly speaking it's not a "distance".
5.4 Role in ML: The Root of Cross-Entropy Loss
Classification problem setup: the true label is a one-hot distribution $p$ (for sample $y=c$, $p(c)=1$, all others 0), and the model outputs a predicted distribution $q = \mathrm{softmax}(z)$. Training aims to "make $q$ as close to $p$ as possible", i.e., minimize
$$ D_{KL}(p | q) = \underbrace{H(p, q)}{\text{cross-entropy}} - \underbrace{H(p)}{\text{constant, independent of the model}} $$
Since the true distribution $p$ is fixed, $H(p)$ is constant, minimizing KL divergence ≡ minimizing cross-entropy. And in the one-hot case, cross-entropy simplifies to "negative log of the predicted probability":
$$ -\log q(y=c) = -\log \hat{p}_c $$
This is why the standard loss for multi-class classification is cross-entropy, and PyTorch/TensorFlow's CrossEntropyLoss is exactly this expression. Revisiting the conclusion from Chapter 4: cross-entropy = maximum likelihood of Bernoulli/categorical distribution. Two paths, one destination — this is precisely why it holds "standard" status.
Why Not MSE for Classification
Applying MSE to classification (with sigmoid/softmax) causes problems: sigmoid saturates at both ends, its derivative approaches 0, and MSE barely provides gradients for saturated-region errors — vanishing gradients, making training extremely slow or even stall. The gradient of cross-entropy is proportional to $\hat{p} - y$ (prediction minus label), naturally avoiding this issue. This isn't magic; it's a hard difference in the mathematical structure of the two loss functions.
5.5 Mutual Information (for awareness)
Mutual information $I(X;Y) = H(X) - H(X\mid Y)$ measures "how much uncertainty about $X$ is reduced by knowing $Y$", a non-linear association measure between distributions. It has a place in feature selection, clustering evaluation, and representation learning (InfoNCE and other self-supervised losses). At the entry level, knowing the concept suffices.
Self-Study Resources · Information Theory
- StatQuest: short video explanations of entropy, cross-entropy, and KL divergence — intuitive and accessible.
- Khan Academy Information Theory course (Brit Cruise series): starting from coding and bits, a painless entry point.
- Claude Shannon's A Mathematical Theory of Communication (1948): the founding paper of information theory, surprisingly readable.
- Cover & Thomas Elements of Information Theory: a systematic textbook, consult as needed.
VI. Optimization: Finding the Parameters
With a loss function (objective) in hand, the next step is finding the parameters that minimize it. That's optimization. The vast majority of ML training follows the same pattern: gradient descent and its variants.
6.1 Intuition of Convex Optimization
Picture the function's values as terrain:
Loss L
│ ∖
│ ∖ · · · ← Non-convex: multiple local minima
│ ∖ ∕
│ V ← Global minimum (convex: a single bowl bottom)
│
└───────────────────────→ parameter w- Convex function: only one bowl bottom (global minimum); any local optimum is the global optimum — linear regression, logistic regression, and SVM objectives are all convex, with theoretical guarantees of "converging to the optimum".
- Non-convex function: the terrain is hilly with multiple ups and downs, offering only local optima — the objectives of deep neural networks are almost always non-convex, but practically we still find good solutions (see §6.4).
The most important intuition about convexity: convex problems let you confidently use gradient descent; non-convex problems rely on engineering tricks (initialization, learning rate schedules, batch normalization) to clear the path.
6.2 Gradient Descent: The Source of All Optimization
Chapter 3 has the tool ready: the negative gradient of loss $L(w)$ points in the steepest direction of descent. So the iterative formula — the most frequently appearing equation in all of ML:
$$ w_{t+1} = w_t - \eta \nabla L(w_t) $$
- $w_t$: parameters at step $t$.
- $\nabla L(w_t)$: gradient of the loss with respect to parameters at the current point.
- $\eta$ (learning rate): step size, the most critical hyperparameter.
The classic visuals of a miscalibrated learning rate:
Loss
│ ↑ Learning rate too large: oscillating or diverging
│ ↕ ↕
│ ↕ ↕
│ ↕ ↕ ↕
│···↕↕···
│ ↑ Learning rate just right: smoothly reaching the bottom
│ ~~~~
│ ~~~~
│~~~~
└────────────────────→ iterations- Learning rate too large: bouncing back and forth past the bottom, or even diverging.
- Learning rate too small: crawling extremely slowly, taking forever to converge.
- Practical solutions: learning rate scheduling (decay), adaptive optimizers like Adam (automatically adjusting step sizes per parameter based on gradient magnitude).
Based on how much data you use per step, there are three variants of gradient descent:
| Variant | Data Used Per Step | Characteristics |
|---|---|---|
| Batch GD | Entire training set | Stable but slow; every step requires a full pass |
| Stochastic GD (SGD) | 1 sample | Fast but noisy; oscillates |
| Mini-batch SGD | A small batch (e.g., 32/64/128) | Practical default: balances stability and speed |
6.3 Lagrange Multipliers and Regularization
Regularization (constraining model complexity) is isomorphic to constrained optimization in optimization. The general form of a constrained optimization problem:
$$ \min_w f(w) \quad \text{s.t.} \quad g(w) \le c $$
Lagrange multiplier method converts it to an unconstrained problem:
$$ \mathcal{L}(w, \lambda) = f(w) + \lambda, (g(w) - c) $$
where $\lambda$ is the multiplier. The key mapping — L2 regularization is simply the Lagrangian form of a parameter-norm constraint:
$$ \underbrace{\min_w ; \text{Loss}(w)}_{\text{original objective}} ;\longleftrightarrow; \underbrace{\min_w ; \text{Loss}(w) + \lambda |w|2^2}{\text{Lagrangian regularized objective}} $$
Intuitive explanation: regularization adds a "surcharge" for "parameters being too large" ($\lambda$ is the penalty tax rate), which is the dual of the constraint optimization "parameters must lie within a sphere of radius $\sqrt{c}$". Larger $\lambda$ means the solution is more constrained — this is the mathematical origin of the "complexity tax". L1 regularization works the same way (replace the sphere with a diamond; at the corners it forces certain coordinates to zero, enabling feature selection). Go deeper in Regularization and Optimization.
6.4 Deep Learning in a Non-Convex World: Why It Still Works
Strictly speaking, the loss function of a deep network is non-convex, and theory only guarantees local optima — yet practical results are surprisingly good. The currently accepted high-dimensional intuition is: when parameter space has extremely high dimensions, local optima are often not "sharp" but rather large near-flat valleys, most of which generalize about equally well; combined with random initialization, mini-batch noise, and learning rate schedules, the optimizer moves like it's traversing rolling hills, always finding its way to somewhere low enough. Understanding this much — "deep learning optimization is engineering-driven, relying on practical tuning" — is sufficient; no need to pursue rigorous proofs. See Optimization for related practical issues.
Self-Study Resources · Optimization
- 3Blue1Brown's neural network series, episode 3: visualizations of gradient descent and backpropagation.
- StatQuest on Gradient Descent: clear plain-language explanation of step size and convergence.
- Andrew Ng's Machine Learning Specialization: the classic introductory treatment of gradient descent and learning rates.
- Boyd & Vandenberghe Convex Optimization: the standard textbook on convex optimization (deep reference, no need to read cover to cover).
VII. Learning Strategy: Learn on Demand, Fill Gaps As You Go
7.1 Three Principles
- Learn on demand: encounter a problem first, then study the math. Encountering "why cross-entropy?" → go read Chapter 5; "why does L2 shrink weights?" → go read §6.3. Learning with questions in mind, applying immediately after — that's when memory sticks best.
- Learn as you go: every time you encounter unfamiliar math in your code, spend 20 minutes that day looking up the relevant section + one instructional video — far more effective than hoarding "I'll study systematically later".
- Intuition before derivations: first understand "what this equation does, what problem it solves"; being able to follow the derivation is enough — no need to independently complete proofs.
7.2 Recommended Learning Sequence
Ordered by "return on investment":
| Phase | What to Learn | Where It's Used | Relevant Section |
|---|---|---|---|
| Week 1 | Vectors, matrix multiplication, dot products | Understanding data and network layers | II |
| Weeks 1–2 | Derivatives, gradients, chain rule | Understanding the training process | III |
| Weeks 2–3 | Distributions, expectation/variance, MLE | Understanding where loss functions come from | IV |
| Week 3 | Entropy, cross-entropy, KL divergence | Understanding classification loss, generative models | V |
| Ongoing | Gradient descent and its variants | Training any model in practice | VI |
| As needed | Eigenvalues, PCA, convexity | Dimensionality reduction, optimization theory | II, VI |
Pacing tip: the first step isn't to finish all the math above, but to first run a complete loop of "data → loss → gradient descent" with the simplest model (linear regression), then go back and fill in the math behind each step one by one. See Learning Paths for a recommended complete entry roadmap.
7.3 Practical Tips
- Visuals first: vectors, matrices, and gradients can all be understood through geometric animations; 3Blue1Brown is the top choice. Write code and plot loss curves across iterations with Matplotlib to intuitively feel convergence.
- Hand-calculate small examples: given a 2×2 matrix and a 2-dimensional sample, compute a forward pass and one gradient update step by hand — more effective than reading a formula ten times.
- Use framework autograd, but stay humble: PyTorch will compute gradients, but you should be able to verbally explain "what this step is computing"; when things go wrong, go back to Chapters 3 and 6.
- Formula reading method: when encountering an unfamiliar formula, first circle the "variables" (are Greek letters parameters or data?), then examine the "structure" (is $\sum$ summing? is $\exp$ normalizing?), and finally ask "what's the intuitive meaning" (is it averaging? measuring distance? penalizing?).
- It's OK to not understand — consult the table: the "concept → use case" table in §2.7, the self-study resources at the end of each chapter, are your portable index. When a term is unclear, flip to the Glossary.
VIII. Further Reading
- What Is Machine Learning — panoramic ML concepts, where math lands
- Learning Paths — recommended complete learning sequence
- Glossary — the "what" dictionary to complement this article
- Model Evaluation and Validation — bias-variance, the statistics behind metrics
- Regularization — engineering implementation of Lagrangian and norm penalties
- Optimization — gradient descent variants and hyperparameter tuning practice
- Deep Learning Fundamentals — where matrix multiplication and the chain rule converge
- Awesome Resources — courses and book maps organized by learning phase
- FAQ — common questions on concepts covered in this article
References
All below are public, freely accessible or formally published, real resources, listed for on-demand reference.
- 3Blue1Brown, Essence of Linear Algebra and Essence of Calculus video series, YouTube (3blue1brown.com).
- 3Blue1Brown, But what is a neural network?, neural network series (including gradient descent and backpropagation), YouTube.
- Gilbert Strang, Introduction to Linear Algebra (5th ed.), Wellesley-Cambridge Press; companion course MIT 18.06, MIT OpenCourseWare (ocw.mit.edu).
- MIT OCW 18.02 Multivariable Calculus and 18.05 Introduction to Probability and Statistics, ocw.mit.edu.
- Joe Blitzstein, Harvard Stat 110 Introduction to Probability, HarvardX/edX and YouTube, freely available.
- Khan Academy, Linear Algebra, Multivariable Calculus, Statistics and Probability, and Information Theory course series, khanacademy.org.
- Josh Starmer, StatQuest channel (gradient descent, entropy, cross-entropy, KL divergence, Bayes, etc.), YouTube.
- Christopher M. Bishop, Pattern Recognition and Machine Learning (PRML), Springer, 2006.
- Ian Goodfellow, Yoshua Bengio, Aaron Courville, Deep Learning ("the Flower Book"), MIT Press, 2016.
- Trevor Hastie, Robert Tibshirani, Jerome Friedman, The Elements of Statistical Learning (ESL), Springer, 2009.
- Claude Shannon, A Mathematical Theory of Communication, Bell System Technical Journal, 1948.
- Thomas Cover, Joy Thomas, Elements of Information Theory (2nd ed.), Wiley, 2006.
- Stephen Boyd, Lieven Vandenberghe, Convex Optimization, Cambridge University Press, 2004.
- Marc Deisenroth, A. Aldo Faisal, Cheng Soon Ong, Mathematics for Machine Learning, Cambridge University Press, 2020 (mml-book.github.io free PDF).
- Andrew Ng, Machine Learning Specialization, Coursera, 2022.