Theme
Linear Models and Logistic Regression
One-line definition: linear models assume that the relationship between target y and features x can be described by a weighted sum of features — linear regression outputs continuous values, logistic regression outputs probabilities. They aren't flashy or "deep," but for any rigorous machine learning project, the first line of modeling code is almost always them.
Flip through the feature engineering sections of Kaggle winning solutions, and you'll find shadows of linear models at the bottom; debug a deep learning model that underperforms XGBoost on tabular data, and the conclusion is often "run a logistic regression baseline first." What this article clarifies is the complete theory and correct usage of this "simplest weapon."
Linear model family overview
┌─────────────────────────────────────────────────────┐
│ Linear regression y ≈ w·x + b Output: continuous (regression) │
│ Logistic regression p(y=1) = σ(w·x+b) Output: probability (binary classification)│
│ Softmax regression p(y=k) ∝ e^(wₖ·x) Output: distribution (multi-class) │
│ Poisson regression E[y] = e^(w·x+b) Output: count (GLM) │
└─────────────────────────────────────────────────────┘1. Two Perspectives: Least Squares and Maximum Likelihood
The same linear model can be derived from two completely different angles. Understanding these two perspectives is the key to understanding the entire supervised learning landscape (the skeleton of supervised learning is also these two languages).
1. The least squares perspective (geometry/optimization)
Suppose we have n samples (xᵢ, yᵢ). Linear regression seeks a line ŷ = w·x + b that minimizes the sum of squared vertical distances from all samples to the line:
Objective: min Σᵢ (yᵢ - (w·xᵢ + b))²
w,bThis is purely an optimization problem: choose a loss function (mean squared error, MSE), then solve for parameters. It doesn't ask how the data was generated or what probabilistic meaning it has — it just geometrically finds the hyperplane "closest to all points." The advantage of this view is directness and computability; the disadvantage is no concept of "confidence": you don't know how reliable this fit is.
2. The maximum likelihood perspective (probability/statistics)
Now reframe the question: assume the true relationship between y and x is linear, but observations are polluted by random noise, i.e.,
yᵢ = w·xᵢ + b + εᵢ , εᵢ ~ N(0, σ²) (noise follows a normal distribution with mean 0)Under this assumption, the probability density of yᵢ is
p(yᵢ | xᵢ) = (1/√(2πσ²)) · exp( -(yᵢ - w·xᵢ - b)² / (2σ²) )Multiply the likelihoods of all samples (assuming independence), take the log, and maximize the log-likelihood — you'll find the objective function is exactly least squares, up to a constant coefficient. In other words:
Key insight
Under the "noise is normally distributed" assumption, maximum likelihood estimation ≡ least squares estimation. Two seemingly different perspectives converge from different paths under the probabilistic assumption. This gives least squares a statistical "credentials": it's not just an arbitrary choice of loss function, but "the most reasonable estimate under a normal noise model."
Conversely, when we swap the loss function from "squared error" to "cross-entropy," we get not linear regression but logistic regression — swapping the loss function = swapping the probabilistic model assumption. This is the main thread for understanding everything below.
3. Four basic assumptions of linear models
Linear regression isn't "force-fit a line onto any data." It has strict applicability conditions (the classic checklist from ISLR Chapter 3):
| Assumption | Meaning | Typical symptoms when violated |
|---|---|---|
| Linearity | y's relationship with each feature xⱼ is approximately linear | Poor fit; residual plot shows curvature |
| Independence | Samples are independent of each other (no autocorrelation) | Common in time series; residuals drift over time |
| Homoscedasticity | Noise variance σ² doesn't vary with x | Residual plot shows a "megaphone" shape |
| Normality | Noise follows a normal distribution (mainly affects inference, not fitting) | Unreliable confidence intervals for small samples |
Note: the first two affect prediction accuracy; the last two mainly affect statistical inference (p-values, confidence intervals). If you only care about prediction accuracy, moderate violations of homoscedasticity and normality are usually tolerable; but if you want causal interpretation like "whether coefficient w is significant," all four need careful checking.
2. Linear Regression: From Formula to Closed-Form Solution
1. Model form and loss
Absorb the bias b into the weights, writing compactly. Let sample matrix X ∈ ℝⁿˣᵈ (n samples, d features), weight vector w ∈ ℝᵈ, objective:
ŷ = X·w (vector form: each sample gets w·xᵢ)
L(w) = (1/n)·‖Xw - y‖² (mean squared error loss)2. Closed-form solution (normal equation)
Least squares is a convex quadratic optimization problem, with an analytical solution: take the gradient of L(w) and set it to zero, yielding
w* = (XᵀX)⁻¹ Xᵀ y (normal equation)This is what LinearRegression().fit(X, y) does internally (actual implementations use numerically more stable QR/SVD decomposition rather than direct inversion). It's done in one step, no iteration needed — this is a luxury unique to linear models: convex + differentiable + quadratic, all three together, and the optimal solution is laid out in the open.
But the cost of the closed-form solution is O(d³) matrix operations (inverting (XᵀX)). When feature dimension d is very large (e.g., millions-dimensional text vectors), or data won't fit in memory, you turn to iterative methods.
3. Gradient descent: when the closed-form solution isn't feasible
Gradient descent is the engine running through all of deep learning (see Optimization and Gradient Descent), taking its cleanest form in linear models:
w ← w - η · (2/n)·Xᵀ(Xw - y)
where η is the learning rate, (2/n)·Xᵀ(Xw - y) is the gradient of the loss w.r.t. wEach iteration takes a small step in the negative gradient direction. Three common variants:
| Variant | How many samples per step | Characteristics |
|---|---|---|
| Batch gradient descent | All n | Most accurate per step, but each step is too expensive for big data |
| Stochastic gradient descent (SGD) | 1 sample | Noisy per step but extremely fast; can escape local pits |
| Mini-batch gradient descent | m (e.g., 32/64) | Compromise, the default choice for deep learning |
4. R²: how to read goodness of fit
R² (coefficient of determination) measures how much of y's variance the model explains:
R² = 1 - SS_res / SS_tot
SS_tot = Σᵢ(yᵢ - ȳ)² total sum of squares (error when predicting with the mean)
SS_res = Σᵢ(yᵢ - ŷᵢ)² residual sum of squares (error when predicting with the model)- R² = 1: perfect fit, zero residuals.
- R² = 0: the model isn't better than "always guessing the mean."
- R² < 0: the model is worse than guessing the mean (usually on the test set; an alert for overfitting or distribution shift).
Three misreadings of R²
- High R² ≠ causality: garbage features can also prop up high R²; good prediction doesn't mean "x causes y."
- R² is in-sample: training-set R² is artificially inflated by default; always report using the test set (or cross-validation).
- Confused with correlation: for simple regression, R² = correlation coefficient squared, but for multiple regression, it's not a simple sum.
A more comprehensive discussion of evaluation metrics is in Model Evaluation and Validation.
3. Logistic Regression: Moving Linear into Probability Space
1. Why classification can't use linear regression directly
Binary labels y ∈ {0, 1}. If you use linear regression to fit, predictions will give meaningless numbers like 1.3, -0.7, outside [0,1]; worse, outliers will pull the line away from the correct direction. Classification doesn't need "continuous values"; it needs "the probability of belonging to class 1" — probabilities must fall within [0, 1].
2. Sigmoid: squishing to (0,1)
Logistic regression's approach: first compute the linear combination z = w·x + b, then squeeze it through the sigmoid function (also called the logistic function):
σ(z) = 1 / (1 + e⁻ᶻ)
p(y=1 | x) = σ(w·x + b) = 1 / (1 + e^-(w·x+b)) σ(z)
1 ──────╴╴╴╴╴
│ ╱
│ ╱ z = 0: σ = 0.5
│╲ This point is the decision boundary
0 ──┼───────→ z
-∞ 0 +∞Three properties of sigmoid worth memorizing:
- Range (0,1): naturally a probability.
- Differentiable everywhere: friendly to backpropagation.
- The more extreme z gets, the more saturated: when z is very large/small, the gradient approaches 0 (this is one source of "vanishing gradients" in deep networks, see Deep Learning Foundations).
3. Cross-entropy: the natural loss for classification
You can't use MSE for classification — the sigmoid + MSE combination makes gradients so small learning nearly stops (in the saturated region). Logistic regression uses cross-entropy / negative log-likelihood as the loss:
For a single sample:
Lᵢ = -[ yᵢ·log(pᵢ) + (1-yᵢ)·log(1-pᵢ) ]
For all samples:
L = -(1/n)·Σᵢ [ yᵢ·log(pᵢ) + (1-yᵢ)·log(1-pᵢ) ]Why is cross-entropy correct? Return to the maximum likelihood perspective: y follows a Bernoulli distribution (single trial of binomial), with probability mass function p^y·(1-p)^(1-y). Take the negative log-likelihood for n independent samples — it's exactly this cross-entropy. Using cross-entropy = assuming a Bernoulli noise model, the same logic as "using MSE = assuming normal noise."
Gradient differences between cross-entropy and MSE:
| Gradient w.r.t. w | Characteristics | |
|---|---|---|
| MSE + sigmoid | Contains σ'(z) factor; in saturation → 0 | Slow convergence, easily stuck |
| Cross-entropy + sigmoid | Proportional to (pᵢ - yᵢ)·xᵢ | More wrong predictions → larger updates; never stalls due to saturation |
Note the gradient is proportional to the error (pᵢ - yᵢ): the more outrageous the prediction, the larger the correction — this is the root of why cross-entropy intuitively "should work well."
4. Decision boundary: linear hyperplane
Taking p = 0.5 as the dividing point, solving σ(w·x+b) = 0.5 gives w·x + b = 0 — a line/hyperplane. This is the meaning of "linear classifier": it draws a line (2D) or hyperplane (high-dimensional) in feature space to split the two classes.
x₂
│ ▲ = class 1
│ ▲ ▲
● │ ▲
● ●│╲ Decision boundary
● │ ╲ w₁x₁+w₂x₂+b=0
● ●● ●│● ╲ ▲
────────┼──────╲───────────→ x₁
│ ╲▲ ▲
│ ╲
│ ● ● ▲
│Two corollaries worth expanding:
- Feature interactions must be built manually: logistic regression's decision boundary is always a line/hyperplane. When faced with an XOR-type distribution (can't be separated by a straight line), it's helpless unless you do feature engineering first to create interaction terms x₁·x₂ (this is precisely one reason why feature engineering exists).
- Probabilities are soft: p=0.6 and p=0.99 are both classified as "class 1," but their confidence levels are worlds apart. In practice, p itself is often used for ranking (e.g., in risk control, sorting by default probability), not just binary classification.
5. Multi-class: softmax regression
Generalizing sigmoid to K classes gives softmax regression (also called multinomial logistic regression):
p(y=k | x) = e^(wₖ·x + bₖ) / Σⱼ e^(wⱼ·x + bⱼ)
Learn a set of weights (wₖ, bₖ) for each class k; output is a probability distribution summing to 1The origin of the name "softmax": e^z exponentiates each logit (score), making large values larger and small values smaller, then normalizes — a "soft" maximum. It is the standard form of every neural network classification head: the last layer of a neural network + softmax is logistic regression's deep version. So logistic regression ≈ "a neural network with no hidden layers," which is the bridge connecting classic models to deep learning.
4. The Role of Regularization: Ridge Regression and Lasso
Linear models have a weakness: when features are many, samples are few, or features are highly correlated, the (XᵀX)⁻¹ in the closed-form solution approaches singularity, w becomes extremely large and unstable — perfect on the training set, a mess on the test set. This is the classic form of overfitting (principles in Overfitting and Regularization).
The idea of regularization is simple and effective: add a penalty term in the loss, forcing w from growing too large.
Ridge regression: L(w) = (1/n)‖Xw - y‖² + λ·‖w‖²₂ (L2, full name "ridge")
Lasso: L(w) = (1/n)‖Xw - y‖² + λ·‖w‖₁ (L1)
Elastic Net: L(w) = (1/n)‖Xw - y‖² + λ₁·‖w‖₁ + λ₂·‖w‖²₂| Method | Penalty | Effect | Characteristics |
|---|---|---|---|
| Ridge regression | Σwⱼ² | Coefficients shrink but don't reach zero | All features retained; first choice for handling collinearity (ill-conditioned matrices) |
| Lasso | Σ|wⱼ| | Coefficients are pressed exactly to zero | Automatic feature selection, sparse solution |
| Elastic Net | Both combined | Has both | More stable than Lasso for high-dimensional strongly correlated features |
Why does L1 make coefficients zero while L2 only shrinks them? Geometric intuition: L2's constraint region is a sphere (fair shrinkage in all directions, only shrinks, never cuts the axes); L1's constraint region is a diamond, and its vertices land exactly on the axes — the optimal solution is more likely to be "pushed" to an axis, meaning some wⱼ exactly equal 0.
L2 constraint (circle) L1 constraint (diamond)
w₂ w₂
│ ╱╲ │ ╲
│╱ ╲ Optimal solution here │ ╲ Optimal solution often lands at a vertex
│ \ ↓ │ ↓ (w₁=0 or w₂=0)
─┼──────────→ w₁ ─┼──────────→ w₁In sklearn, logistic regression's regularization parameter is called C (note it's the inverse: C = 1/λ, smaller C means stronger regularization; needs hyperparameter tuning). For linear regression's ridge/Lasso, use alpha (alpha = λ, larger alpha means stronger regularization). In practice, almost always keep regularization on — even the mild default of C=1.0 is better than defenseless.
5. Generalized Linear Models: Making "Linear" a Family
Linear regression handles continuous values, logistic regression handles 0/1, but what about "website click counts" (non-negative integers), "loan amounts" (right-skewed positive numbers)? The answer is to unify linear regression and logistic regression into a larger framework — Generalized Linear Models (GLM).
1. Three elements of GLM
Nelder and Wedderburn proposed GLM in 1972, composed of three components:
1. Random component: y follows some "exponential family" distribution (Gaussian, Bernoulli, Poisson, Gamma, ...)
2. Linear predictor: η = w·x + b (always linear; this is why the model is named "linear")
3. Link function g: E[y] = g⁻¹(η), connecting the mean to the linear combinationThe key leap: allow y's distribution to be non-normal, and allow E[y]'s relationship with x to be nonlinear (through the link function), but inside there's still a linear expression η = w·x + b — so the estimation framework (weighted least squares, iteratively reweighted least squares, convex optimization) can all be reused.
2. Classic members overview
| Model | y's distribution | Link function g(μ) | Suitable data |
|---|---|---|---|
| Linear regression | Gaussian N(μ, σ²) | Identity: μ | Continuous values |
| Logistic regression | Bernoulli Ber(p) | logit: log(p/(1-p)) | 0/1 labels |
| Poisson regression | Poisson Poi(λ) | log: log(λ) | Counts (clicks, accidents) |
| Gamma regression | Gamma distribution | inverse: 1/μ | Positive right-skewed (premiums, duration) |
3. Poisson regression example
Predict "daily click count for a webpage": counts are non-negative integers, usually right-skewed; linear regression would predict negative numbers. Poisson regression assumes y ~ Poisson(λ), using a log link:
log E[y] = w·x + b ⟹ E[y] = e^(w·x + b)
For samples: yᵢ ~ Poisson( λᵢ = e^(w·xᵢ + b) )The log link guarantees predicted counts are always non-negative; when λ is large, Poisson approaches normal, also explaining why Gaussian models can suffice for large counts. In practice, count data often has "overdispersion" (variance > mean), at which point you can upgrade to negative binomial regression — another "change distribution, change link" in the GLM framework.
GLM's engineering value: the same fit/predict interface, swap the distribution and link function to serve different business scenarios. sklearn's sklearn.linear_model provides PoissonRegressor, GammaRegressor, TweedieRegressor, and other GLM implementations.
6. When to Use Linear Models
Linear models aren't cool, but they have an entire set of positions that no one can replace.
1. Baseline value: let simple models hold the bottom
The first step of any modeling project should be "run a logistic regression/linear regression" (with standardization and regularization). The reasoning is solid:
- Fast: minute-level training; get data pipelines and evaluation scripts running first.
- Set a benchmark: complex models (trees, neural networks) must prove they significantly exceed this baseline; otherwise they're trading complexity for nothing.
- Validate data: if a linear model performs terribly, it usually means data leakage, features not matching labels, or abnormal target distribution — expose problems first, then talk about advanced models.
This is a repeatedly verified engineering rule: a baseline isn't an optional step; it's the "passing line" for complex models (there's a dedicated entry in Common Pitfalls for "skipping the baseline and going straight to complex models").
2. Interpretability: the whitest of white boxes
A linear model's interpretability is priceless:
Coefficient wⱼ's meaning: holding all other features constant, for every 1-unit increase in xⱼ, predict y changes by wⱼ on average
Example: House price = 5K/m² × area + 80K × is school-district + 300K
→ "for each additional 1m² of area, house price increases on average by 5000 yuan"Coefficients can be directly explained to business stakeholders, audited, questioned, and turned into regulations. In highly regulated scenarios that require explaining decision rationale — financial risk control, healthcare, judiciary — linear models are a powerful tool for compliance. Tree models' SHAP explanations (see Tree Models and Ensemble Learning) can approach this interpretability, but they can't achieve the conciseness of "one line of coefficients goes everywhere."
3. Comparison with tree models and deep learning
| Dimension | Linear models | Tree models (RF/GBDT) | Deep learning |
|---|---|---|---|
| Nonlinearity / interactions | Need manual feature engineering | Learned automatically (axis-aligned splits) | Learned automatically (arbitrary function approximation) |
| Tabular data | Baseline, often sufficient | Usually the strongest | Usually underperforms tree models |
| Images / text / audio | Barely used | Barely used | Absolute home field |
| Sample size requirement | Low (can start with hundreds) | Medium | High (shines only with 10K+) |
| Interpretability | ★★★ | ★★ (SHAP-assisted) | ★ (black box) |
| Training cost | Seconds~minutes | Minutes~hours | GPU-hours~days |
The empirical rule (repeatedly verified by competitions):
Tabular data + medium-small sample → start with logistic regression, then tree models (GBDT family)
Tabular data + extremely large sample → tree models still stable; deep learning is competitive but needs heavy tuning
Unstructured data (image/text/audio) → deep learning; linear models only as the final classification headLinear models aren't "outdated technology being phased out" on this spectrum; they are complementary underlying components alongside tree models and deep learning — the softmax output layer of a neural network is logistic regression, and that layer is often the only "interpretable part" of the model that can be precisely explained.
7. Practice: Complete Binary Classification with Logistic Regression
Using sklearn's built-in Wisconsin Breast Cancer dataset (569 samples, 30 cell nucleus features, binary classification: malignant/benign) as an example, walk through the complete flow: split → standardize → train → evaluate.
Why must you standardize?
Logistic regression is sensitive to feature scales — regularization penalty Σwⱼ² would "unfairly" penalize large-scale features, and gradient descent converges extremely slowly when scales are uneven. StandardScaler normalizes each feature to mean 0, variance 1. Fit only on the training set, before splitting; never fit on the full dataset — this is the most common rookie mistake that leaks test information into training.
python
import numpy as np
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, precision_score, recall_score, roc_auc_score, confusion_matrix
# ── ① Data ─────────────────────────────────────────────
data = load_breast_cancer()
X, y = data.data, data.target # y: 0=malignant, 1=benign
# ── ② Split train/test (stratified sampling, keep class proportions) ──
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, random_state=42, stratify=y)
# ── ③ Standardize: only fit on the training set! ────────────────────────────
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test) # only transform, don't re-fit
# ── ④ Train logistic regression (C is the inverse of regularization strength) ──
model = LogisticRegression(C=1.0, max_iter=2000, random_state=42)
model.fit(X_train_scaled, y_train)
# ── ⑤ Predict and evaluate ───────────────────────────────────────
y_pred = model.predict(X_test_scaled)
y_proba = model.predict_proba(X_test_scaled)[:, 1] # probability of class 1
print(f"Accuracy : {accuracy_score(y_test, y_pred):.4f}")
print(f"Precision : {precision_score(y_test, y_pred):.4f}")
print(f"Recall : {recall_score(y_test, y_pred):.4f}")
print(f"AUC : {roc_auc_score(y_test, y_proba):.4f}")
print("Confusion matrix:")
print(confusion_matrix(y_test, y_pred))Output (typical result, reproducible with the random seed):
Accuracy : 0.9825
Precision : 0.9901
Recall : 0.9804
AUC : 0.9976
Confusion matrix:
[[ 60 3]
[ 0 108]]How to read the results
- AUC ≈ 0.998: the model's ranking ability for "who looks more benign" is extremely strong — even if the decision threshold drifts, the ranking holds.
- 3 malignant samples were misclassified as benign in the confusion matrix: in this scenario where "missing a malignant case" is far more costly than "falsely flagging benign," you should lower the decision threshold (e.g., classify as malignant if p ≥ 0.3), sacrificing a bit of precision for recall. The threshold is always a business decision, not a model parameter.
- For more sensitive metrics, return to Model Evaluation and Validation for the full discussion of metric matrices.
Coefficient interpretation and feature importance
python
# Rank features by absolute coefficient value, see what the model depends on most
coef = model.coef_[0]
feat_names = data.feature_names
ranked = sorted(zip(feat_names, coef), key=lambda t: -abs(t[1]))
for name, c in ranked[:5]:
print(f"{name:28s} coefficient = {c:+.3f}")Note: only compare coefficients after standardization (all features on the same scale); comparing absolute values of unstandardized coefficients is wrong — another pitfall hiding in implementation details.
8. Tradeoffs and Decision Points
1. The linearity assumption: convenience vs. distortion
Linear models' core cost is the "features and targets are approximately linear and additive" assumption. In reality, y and x are often nonlinear and interactive. Two paths out:
- Feature engineering: move nonlinearity into features (polynomial, log, binning, cross-terms), making the linear model "still linear in the transformed space" — this is the most important craft in traditional machine learning, see Feature Engineering.
- Switch models: if data is too nonlinear and interactions too complex, upgrade to tree models or deep learning, letting the model learn automatically (at the cost of interpretability and more tuning).
A practical criterion: when dimension d is much smaller than sample n, and features have been seriously preprocessed, linear models often perform nearly as well as complex models; otherwise, upgrade sooner rather than later.
2. Five disciplines when using linear models
- Standardize first: scale differences distort regularization and convergence (see Section 7).
- Handle feature collinearity: highly correlated features make coefficient signs jump and explanations distorted — absorb it with ridge regression, or use Lasso / variance inflation factor (VIF) to eliminate.
- Keep regularization on: the default C=1.0 is mild but better than defenseless; treat "preventing overfitting" as the default config, not a post-hoc fix.
- Evaluate on the test set / AUC: training-set R² and accuracy are artificially inflated by default; always speak with cross-validation or holdout test sets.
- Coefficients ≠ causality: no matter how strong the correlation, it's just a prediction tool; business attribution requires experimental design (A/B tests, randomization).
3. When you must "upgrade"
When you see the following signals, the linear model's plate can no longer hold:
• Residual plots / partial dependence plots show obvious nonlinear relationships that feature engineering can't fix
• Strong interactions between features that make manually constructed cross-terms unmaintainable
• Data is unstructured input (images/text/audio)
• Sample size is massive and compute is sufficient, and deep learning can squeeze out a few more pointsBut remember: the direction of upgrading is climbing step by step from the baseline: logistic regression → logistic regression with feature engineering → GBDT → deep learning. Compare each step against the baseline and confirm the gains are worth the complexity. This "baseline-driven" methodology is more fully developed in Tree Models and Ensemble Learning and Model Evaluation and Validation.
Further Reading
- What is Machine Learning — Linear models' position in the ML landscape
- Supervised Learning — Formal framework for classification/regression; the context of this article's two perspectives
- Model Evaluation and Validation — Accuracy, confusion matrix, AUC, and bias-variance tradeoff
- Overfitting and Regularization — Complete principles of ridge/Lasso; the theoretical foundation of this article's Section 4
- Feature Engineering — Linear models' nonlinearity and interactions are supplemented by feature engineering
- Optimization and Gradient Descent — Beyond closed-form solutions, gradient descent is the common engine of linear models and deep learning
- Tree Models and Ensemble Learning — The next upgrade step for tabular data
- Deep Learning Foundations — How logistic regression evolves into neural networks
- Unsupervised Learning — A clustering perspective outside the linear family
- ML vs AI Boundary — Linear models' relationship with the "AI" boundary
- Common Pitfalls — Standardization leakage, data leakage, and other frequent failure points
- Framework Comparison — sklearn's ecosystem position relative to other libraries
- Math Primer — Quick reference for matrices, probability, and optimization
- Glossary — Quick reference for sigmoid, MSE, AUC, and other terms
References
- James, Witten, Hastie, Tibshirani. An Introduction to Statistical Learning (ISLR, Springer, 2nd ed. 2021) — Standard textbook on linear regression (Chapter 3), classification/logistic regression (Chapter 4), and regularization (Chapter 6)
- Hastie, Tibshirani, Friedman. The Elements of Statistical Learning (ESL, 2nd ed. 2009) — Mathematical derivations for linear models and GLM (Chapters 3, 4); the source of this article's closed-form and gradient formulas
- scikit-learn User Guide: Linear models — Official docs for ridge, Lasso, logistic regression, and GLM
- scikit-learn API docs: LogisticRegression — Parameter explanations for
C,max_iter,predict_probaused in this practice - scikit-learn API docs: Wisconsin breast cancer dataset — Official documentation for the dataset used in Section 7
- Andrew Ng. CS229 Machine Learning Lecture Notes (Stanford) — Maximum likelihood derivations for linear regression and logistic regression (Chapters 1-2)
- Goodfellow, Bengio, Courville. Deep Learning (the "Deep Learning Book") — In-depth discussion of softmax and cross-entropy (Chapter 6)