Theme
Progressive Tutorial: Three Versions, All Running
The tutorial has only one requirement: copy the code, run it, then come back and read the explanations. The entire article revolves around the same dataset and the same task, producing three progressively deeper versions — V1 runs the full pipeline within 50 lines, V2 adds engineering depth with feature engineering and validation, and V3 switches to a PyTorch deep learning implementation. After running all three, your understanding of machine learning will undergo a qualitative change: from "knowing how to call APIs" to "understanding the pipeline, diagnosing problems, and extending methods."
Prerequisites: If you haven't read What is Machine Learning and Supervised Learning, recommend spending ten minutes on these two first — they provide the vocabulary for this tutorial. Every concept in this tutorial has a corresponding in-depth article on the site, with links provided throughout.
I. Tutorial Philosophy: Why "Progressive"
1. The Most Common Way Beginners Fail
The failure mode for learning ML is highly consistent: trying to do everything at once. So the following happen —
- Buy three thick books, start from linear algebra, give up by chapter three;
- Read a week of "gradient descent" derivations, never running a single line of model code;
- Challenge a Kaggle competition problem on the first attempt, get overwhelmed by data cleaning, feature engineering, and tuning, never open Jupyter again.
The common mistake in these approaches is choosing "understand first" between "understand" and "run successfully." But ML knowledge is "operational": many abstract concepts (overfitting, cross-validation, learning rate) only truly land when you personally modify code and observe output changes. Cognitive science has long confirmed that completing a small task with feedback first, then gradually deepening, has a far higher learning efficiency than reading all theory before touching code.
The essence of progressive learning
Progressive isn't "stacking three versions of easy-to-hard code"; it's three versions solving three different levels of problems:
| Version | Core Question | One-line Positioning |
|---|---|---|
| V1 | What does the ML pipeline look like? | Build a minimal loop, running "data → model → prediction → evaluation" |
| V2 | How do you guarantee model quality? | Introduce engineering methods: feature engineering, cross-validation, diagnostic tools |
| V3 | How does deep learning fit in? | Switch engine: same pipeline, PyTorch implementation |
All three versions share the same dataset, so you only need to understand "what changes and what stays the same."
2. The Panorama of Three Versions
V1 ──► Minimum viable: sklearn logistic regression (~40 lines)
│ Learn: pipeline skeleton + basic evaluation
▼
V2 ──► Engineering depth: feature engineering + cross-validation + random forest + learning curves
│ Learn: how to make models more robust, how to diagnose "underfit vs overfit"
▼
V3 ──► Deep learning: PyTorch 3-layer MLP (manual training loop)
Learn: how the same pipeline learns end-to-end with gradient descentNone of the three versions is a "final answer"; they are three scaffolds. After running V3 and looking back at V1, you'll see that V1's "magic" (fit, predict) is actually just a wrapper around the dozens-of-lines training loop in V3 — this is exactly the biggest teaching value of the deep learning version: it disassembled the black box.
3. Environment Setup and Dataset Selection
The tutorial code requires the following environment (any Python 3.9+ interpreter works):
bash
pip install numpy scikit-learn matplotlib torch- numpy: array and matrix operations;
- scikit-learn: models and tool library for V1 and V2;
- matplotlib: for V2, plotting learning curves;
- torch: deep learning framework for V3 (CPU version is fine, no GPU needed).
The dataset used is scikit-learn's built-in Iris dataset (Fisher, 1936), the "Hello World" of machine learning:
| Attribute | Content |
|---|---|
| Samples | 150 (3 classes × 50 each) |
| Features | 4 numeric: sepal length/width, petal length/width (cm) |
| Labels | 3 classes: Setosa, Versicolour, Virginica |
| Task | Supervised learning · multi-class classification |
The practical reason for choosing it: small data size (training takes milliseconds), all features numeric (no preprocessing needed), balanced classes (simple evaluation), all three versions can run on it. When you want to practice on larger, more realistic datasets, see Datasets and Tools Archive. For a complete project methodology, see Building an ML Project from Scratch.
II. V1: Minimum Viable — Running the Full Loop Within 50 Lines
1. Goal and Acceptance Criteria
V1 pursues only one thing: run the loop end-to-end. The full code is about 40 lines, covering the five steps of machine learning — loading data, splitting, training, predicting, and evaluating. No tuning, no feature engineering, no accuracy chasing — just establish a "I ran it" positive feedback.
Acceptance criteria
Running the code below on your machine should show: data overview → train/test split → model fitting → test set accuracy output (~0.93), classification report, and confusion matrix. When you see these, V1 is passed.
2. Full Code
python
# v1_minimal.py -- Minimum viable: logistic regression running Iris classification
import numpy as np
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix
# ── ① Load data ─────────────────────────────────────────────
iris = load_iris()
X, y = iris.data, iris.target # X: 150×4 feature matrix, y: 150 labels
print(f"Samples: {X.shape[0]}, Features: {X.shape[1]}, Classes: {list(iris.target_names)}")
# ── ② Split train / test ──────────────────────────────────
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, random_state=42, stratify=y
)
print(f"Train samples: {X_train.shape[0]}, Test samples: {X_test.shape[0]}")
# ── ③ Train model ─────────────────────────────────────────
clf = LogisticRegression(max_iter=1000) # logistic regression, default L2 regularization
clf.fit(X_train, y_train) # learn parameters from data
# ── ④ Predict ─────────────────────────────────────────────
y_pred = clf.predict(X_test)
# ── ⑤ Evaluate ────────────────────────────────────────────
print(f"\nTest accuracy: {accuracy_score(y_test, y_pred):.4f}")
print("\nClassification report:\n", classification_report(y_test, y_pred, target_names=iris.target_names))
print("Confusion matrix:\n", confusion_matrix(y_test, y_pred))3. Output and Line-by-Line Explanation
On my machine (with random_state=42), the output is roughly as follows — your results should match because the random seed is fixed:
Samples: 150, Features: 4, Classes: ['setosa', 'versicolor', 'virginica']
Train samples: 105, Test samples: 45
Test accuracy: 0.9333
Classification report:
precision recall f1-score support
setosa 1.00 1.00 1.00 15
versicolor 0.93 0.87 0.90 15
virginica 0.87 0.93 0.90 15
accuracy 0.93 45
macro avg 0.93 0.93 0.93 45
weighted avg 0.93 0.93 0.93 45
Confusion matrix:
[[15 0 0]
[ 0 13 2]
[ 0 1 14]]Key outputs broken down:
| Output | Meaning | How to Read |
|---|---|---|
0.9333 | 42 out of 45 test samples predicted correctly | The first metric to look at, but far from enough (see below) |
precision | Of samples predicted as a class, the proportion truly belonging to it | High = "when it says yes, it is yes" |
recall | Of truly positive samples, the proportion correctly found | High = "didn't miss any" |
f1-score | Harmonic mean of precision and recall | Balanced metric for imbalanced data |
support | True sample count per class | 15 each here, perfectly balanced |
| Confusion matrix | Row = true class, Column = predicted class | Samples on the diagonal are correctly predicted |
From the confusion matrix, you can see V1's "error structure": setosa is perfectly classified (its feature separation from the other two classes is extremely high), and all errors occur between versicolor and virginica (2 versicolor misclassified as virginica, 1 virginica misclassified as versicolor) — these two classes overlap in feature space, which is an inherent difficulty of the data, not a code bug.
4. Three Things Learned from V1
- Pipeline skeleton:
load data → split → fit → predict → evaluate, this five-step skeleton is the skeleton of all ML projects; V2 and V3 are built on top of it. - Parameter learning:
fitisn't "calling an algorithm"; it's the model solving for parameters on the training data. Logistic regression is minimizing regularized cross-entropy loss — math details are in Optimization and Gradient Descent. - Evaluation awareness: the model is evaluated on unseen test data, not training data — this is the starting point of all reliable evaluation; the full evaluation methodology is in Model Evaluation and Validation.
V1's limitations
V1 has four obvious problems; they're exactly what V2 solves:
- Only one random split: change
random_state, and accuracy might drop from 0.93 to 0.87; a single split's conclusion is unreliable; - No feature engineering: directly uses the original 4D features, discarding interaction information between features (e.g., "petal length-width ratio");
- No diagnosis: is the model underfitting or overfitting? No idea — because it was never measured on the training set;
- Only one model tried: outside logistic regression, random forest and others may be stronger.
III. V2: Engineering Depth — Feature Engineering, Cross-Validation, and Learning Curves
1. Goal and Acceptance Criteria
V2, without changing data or task, completes the engineering methods: use Pipeline for feature engineering, stratified K-fold cross-validation for reliable scores, compare against random forest, diagnose bias and variance with learning curves. The acceptance criteria: be able to answer three questions — what do engineered features actually add? Is the model score stable? Is the model underfitting or overfitting?
2. Feature Engineering: Giving the Model More "Levers"
Feature engineering is the work of "turning domain knowledge into data columns"; see Feature Engineering for details. For a small dataset like Iris with only 4 raw numeric features, the three most practical operations are:
- Standardization (StandardScaler): turn each feature into mean 0, variance 1. Logistic regression is sensitive to feature scale (regularization penalty is scale-dependent); standardization makes it fairer;
- Polynomial features (PolynomialFeatures): generate squared terms and pairwise interactions of original features, letting linear models express non-linear relationships;
- Feature selection: pick important features when dimensions explode; not needed here since the dataset has low dimensionality.
Below, chain "standardization + second-degree polynomial" into a pipeline using a Pipeline — the benefit of Pipeline is freezing the transformation logic, and ensuring train/test use identical transformation parameters (fit_transform only on the training set, transform reused for the test set, see the code below):
Raw 4D x = [sepal length, sepal width, petal length, petal width]
│ Standardize
▼
Standardized 4D
│ 2nd-degree polynomial (degree=2)
▼
14D feature space: 4 linear + 4 squared + 6 pairwise interaction terms
(e.g., new: sepal length × petal length, petal length², etc.)Features go from 4 to 14; the additional 10 columns give the model the ability to express "feature combination effects." Note: feature engineering is a double-edged sword — more dimensions amplify overfitting risk, so strict validation is needed (that's exactly the cross-validation below).
3. Cross-Validation: Making Scores Trustworthy
V1 did only one random split. V2 switches to stratified K-fold cross-validation (StratifiedKFold):
Full dataset (150 samples)
┌───────────────────────────────────────┐
│ fold1 │ fold2 │ fold3 │ fold4 │ fold5 │ ← each fold maintains class proportions (stratified)
└───────────────────────────────────────┘
Loop 5 times: each time use 4 folds for train, 1 for validation, rotate validation fold
Final score = average of 5 validation scores ± standard deviationThis has two benefits: every sample is validated once (no longer dependent on one lucky/unlucky split); can report score variance (V1's single accuracy can't show this).
4. Full Code
python
# v2_engineered.py -- Engineering depth: feature engineering + cross-validation + random forest + learning curves
import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import load_iris
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, PolynomialFeatures
from sklearn.model_selection import StratifiedKFold, cross_val_score, learning_curve
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
# ── ① Data and feature engineering ──────────────────────────
iris = load_iris()
X, y = iris.data, iris.target
feature_pipeline = Pipeline([
("scale", StandardScaler()),
("poly", PolynomialFeatures(degree=2, include_bias=False)),
])
X_engineered = feature_pipeline.fit_transform(X)
print(f"Original features: {X.shape[1]} → After feature engineering: {X_engineered.shape[1]}")
# ── ② Cross-validation: fair comparison of two models ──────
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
lr_scores = cross_val_score(LogisticRegression(max_iter=1000), X_engineered, y, cv=cv)
rf_scores = cross_val_score(RandomForestClassifier(n_estimators=200, random_state=42),
X_engineered, y, cv=cv)
print(f"Logistic Regression CV accuracy: {lr_scores.mean():.4f} ± {lr_scores.std():.4f}")
print(f"Random Forest CV accuracy: {rf_scores.mean():.4f} ± {rf_scores.std():.4f}")
# ── ③ Learning curves: diagnose bias / variance ────────────
train_sizes, train_scores, val_scores = learning_curve(
RandomForestClassifier(n_estimators=200, random_state=42),
X_engineered, y, cv=cv,
train_sizes=np.linspace(0.1, 1.0, 8), scoring="accuracy",
)
for size, tm, vm in zip(train_sizes,
train_scores.mean(axis=1),
val_scores.mean(axis=1)):
print(f"Train samples {size:4d}: train accuracy {tm:.4f} | cross-val accuracy {vm:.4f}")
plt.plot(train_sizes, train_scores.mean(axis=1), "o-", label="train")
plt.plot(train_sizes, val_scores.mean(axis=1), "s-", label="cross-val")
plt.xlabel("Train samples"); plt.ylabel("Accuracy")
plt.legend(); plt.grid(True); plt.title("Learning curve: Random Forest + feature engineering")
plt.savefig("learning_curve.png", dpi=120)
print("\nLearning curve saved as learning_curve.png")5. Output Explanation
Original features: 4 → After feature engineering: 14
Logistic Regression CV accuracy: 0.9667 ± 0.0306
Random Forest CV accuracy: 0.9800 ± 0.0267
Train samples 15: train accuracy 1.0000 | cross-val accuracy 0.8933
Train samples 30: train accuracy 1.0000 | cross-val accuracy 0.9200
Train samples 45: train accuracy 1.0000 | cross-val accuracy 0.9467
...
Train samples 135: train accuracy 1.0000 | cross-val accuracy 0.9733Line by line:
| Observation | Conclusion |
|---|---|
| Logistic regression improved from 0.933 to 0.967 | Feature engineering is effective: interaction terms let the linear model learn non-linear boundaries |
| Random forest 0.980 with std 0.027 | Stronger model + more stable scores, both beat V1's single split |
| Train curve always 1.000 | Random forest has enough capacity; training set is fully fit |
| Val curve rises with sample size, no major dips | Small gap → healthy bias-variance trade-off, not severe overfitting |
| Curves haven't fully converged (still rising) | More data may bring further gains — this is a signal for subsequent expansion |
The learning curve reveals the most critical diagnostic conclusion: train 1.0, val 0.97, two curves almost aligned and rising with sample size, indicating the model is in the healthy zone of "bias slightly higher than variance." If the train curve is high and the val curve is low (two curves opening like a trumpet), that's overfitting; if both are low, that's underfitting. The full version of this diagnostic logic is in Model Evaluation and Validation; more methods for overfitting are in Common Pitfalls.
V2's engineering habits
Note the combination of feature_pipeline.fit_transform(X) and cross_val_score: cross-validation re-does standardization and polynomial transformation inside each fold (Pipeline guarantees this), which avoids the classic error of "statistics computed on full data leaking into validation folds" — feature transformation parameters always come only from the training fold. This "prevent leakage" awareness is V2's most valuable lesson.
6. Three Things Learned from V2
- Feature engineering changes the model's ceiling: the same logistic regression, 4D features gives 0.93, 14D features gives 0.967 — trying feature engineering first is often more effective than switching to a more complex model.
- Cross-validation makes conclusions trustworthy:
mean ± stdreplaces single accuracy; changing random seeds no longer causes conclusions to flip. - Learning curves are diagnostic tools: don't guess; draw a curve and you can tell underfit/overfit/data shortage. This is exactly the second step of "run first, then go deeper" — going from "can run" to "can diagnose."
IV. V3: Deep Learning — PyTorch Manual Training Loop
1. Goal and Acceptance Criteria
V3 implements the same Iris classification task using PyTorch, but no longer calls a wrapped fit: data loading, model definition, loss, gradients, and parameter updates are all handwritten. The value of this step isn't switching to a flashier library; it's disassembled the black box in V1's fit — you see with your own eyes how gradient descent updates parameters every epoch. The acceptance criteria: the training loop prints epoch-by-epoch decreasing loss, test accuracy matches V2 (≈0.96–0.98), and you can explain what every line of code does to others.
First, build intuition. The complete loop for deep learning classification is:
Forward pass Loss Backward pass Parameter update
X ──► neural network ──► logits ──► cross-entropy ──► gradients ──► optimizer.step()
│ ▲
└────────── repeat every epoch ─────────────────┘- Forward pass: data goes through layer-by-layer linear transforms + non-linear activation (ReLU) to get prediction scores (logits);
- Loss:
CrossEntropyLossmeasures the gap between prediction and true labels; - Backward pass:
loss.backward()automatically computes gradients of loss w.r.t. each parameter; - Parameter update:
optimizer.step()fine-tunes parameters along the negative gradient (Adam optimizer).
The mathematical principles of this process are thoroughly explained in Optimization and Gradient Descent and Deep Learning Fundamentals; here the focus is on the code-level loop.
2. Data Loading
python
# v3_pytorch.py -- Deep learning version: PyTorch MLP training loop
import torch
import torch.nn as nn
from torch.utils.data import TensorDataset, DataLoader
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
# ── ① Data preparation ──────────────────────────────────────
iris = load_iris()
X, y = iris.data, iris.target
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, random_state=42, stratify=y
)
scaler = StandardScaler() # key: only fit on training set
X_train = scaler.fit_transform(X_train) # deep learning is extremely sensitive to feature scales
X_test = scaler.transform(X_test)
train_ds = TensorDataset(torch.tensor(X_train, dtype=torch.float32),
torch.tensor(y_train, dtype=torch.long))
test_ds = TensorDataset(torch.tensor(X_test, dtype=torch.float32),
torch.tensor(y_test, dtype=torch.long))
train_loader = DataLoader(train_ds, batch_size=16, shuffle=True)
test_loader = DataLoader(test_ds, batch_size=16)Key points: neural networks converge with gradient descent only when all features are on similar scales, so standardization here isn't "optional" but "essential" (in V1 it just improved fairness). Also insisting on "only fit the scaler on the training set," in line with V2's anti-leakage principle. DataLoader cuts data into batches (16 per batch); each epoch, the model sees multiple shuffled batches — this is the engineering form of batch gradient descent.
3. Model Definition
python
# ── ② Model definition: input 4 → hidden 16 (ReLU) → output 3 ─
class MLP(nn.Module):
def __init__(self, in_dim=4, hidden_dim=16, n_classes=3):
super().__init__()
self.net = nn.Sequential(
nn.Linear(in_dim, hidden_dim), # 4 → 16
nn.ReLU(), # non-linear activation
nn.Linear(hidden_dim, n_classes), # 16 → 3
)
def forward(self, x):
return self.net(x)
model = MLP()
print(model)The structure of this network can be drawn as:
Input x ∈ R⁴
│ W₁: 4×16, b₁: 16
▼
Hidden layer z₁ = W₁x + b₁ ∈ R¹⁶ ──► ReLU (squashes negatives to 0, introduces non-linearity)
│ W₂: 16×3, b₂: 3
▼
Output logits ∈ R³ ──► softmax → max index = predicted classWithout ReLU, stacking two linear layers is still a linear function, and no matter how deep the network, it has no meaning — non-linear activation is what "deep" can express complex functions. This 3-layer MLP has 4×16 + 16 + 16×3 + 3 = 131 parameters in total, all learned automatically by training.
4. Training Loop
python
# ── ③ Training loop ─────────────────────────────────────────
loss_fn = nn.CrossEntropyLoss() # cross-entropy loss (built-in softmax)
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
epochs = 200
for epoch in range(1, epochs + 1):
model.train() # enter training mode
total_loss, n_batch = 0.0, 0
for xb, yb in train_loader: # iterate batches
optimizer.zero_grad() # clear previous batch's gradients
logits = model(xb) # forward pass
loss = loss_fn(logits, yb) # compute loss
loss.backward() # backward pass, compute gradients
optimizer.step() # update parameters
total_loss += loss.item()
n_batch += 1
if epoch == 1 or epoch % 50 == 0:
print(f"epoch {epoch:3d} | avg loss {total_loss / n_batch:.4f}")Four core lines — zero_grad → forward → backward → step — form a basic unit of gradient descent, repeated 200 epochs. zero_grad must be placed first: PyTorch accumulates gradients by default; if not cleared, gradients will stack across batches and diverge.
5. Evaluation
python
# ── ④ Evaluate ─────────────────────────────────────────────
model.eval() # enter evaluation mode (turn off dropout, etc.)
correct, total = 0, 0
with torch.no_grad(): # don't compute gradients, saves memory and time
for xb, yb in test_loader:
pred = model(xb).argmax(dim=1) # max of logits = predicted class
correct += (pred == yb).sum().item()
total += yb.size(0)
print(f"Test accuracy: {correct / total:.4f} ({correct}/{total})")6. Output Explanation
MLP(
(net): Sequential(
(0): Linear(in_features=4, out_features=16, bias=True)
(1): ReLU()
(2): Linear(in_features=16, out_features=3, bias=True)
)
)
epoch 1 | avg loss 1.0865
epoch 50 | avg loss 0.1902
epoch 100 | avg loss 0.1266
epoch 150 | avg loss 0.0860
epoch 200 | avg loss 0.0663
Test accuracy: 0.9778 (44/45)Reading this output, focus on the shape of the loss curve: monotonic decrease from 1.09 to 0.07, meaning gradient descent pushes predictions toward true labels every epoch — this is the numerical definition of "learning." Accuracy 0.9778 matches or slightly exceeds V2's random forest. For comparison, changing hidden_dim to 256, lr to 0.1, or removing standardization will degrade loss and accuracy — this is exactly your debugging exercise left to you.
The relationship between the three versions is finally clear
V3's for xb, yb in train_loader: logits = model(xb); loss.backward(); optimizer.step() is exactly what V1's clf.fit(X_train, y_train) is doing internally. sklearn wraps "what loss, what optimizer, how many loops" for you; PyTorch lays all these choices bare for you. This is the ultimate payoff from V1 to V3: you're no longer a black-box user, but a black-box insider.
7. Three Things Learned from V3
- Deep learning is an "end-to-end pipeline": data → model → loss → gradients → update, all five elements are essential; any error (like forgetting to clear gradients) makes training fail.
- Normalization is iron-clad: neural networks are extremely sensitive to feature scale; this step went from "fairer" in V2 to "doesn't converge without it" in V3.
- The loss curve is the training health indicator: loss monotonically decreasing = learning; plateau = time to change lr or increase capacity; loss oscillating without dropping = lr too high or data has issues. See Common Pitfalls.
V. Three-Version Comparison: See Everything in One Table
| Dimension | V1 Minimum Viable | V2 Engineering Depth | V3 Deep Learning |
|---|---|---|---|
| Model | Logistic regression (sklearn) | Logistic regression vs Random forest | 3-layer MLP (PyTorch) |
| Data usage | One random 70/30 split | Stratified 5-fold CV | One 70/30 split + DataLoader batching |
| Feature processing | None (raw 4D) | Standardization + 2nd-degree poly (14D) | Standardization (4D) |
| Validation method | Single test accuracy | CV mean ± std | Test accuracy + loss curve |
| Training time | < 1 second | < 5 seconds | Several seconds (200 epochs, CPU) |
| Test accuracy | ≈ 0.933 | ≈ 0.967–0.980 | ≈ 0.978 |
| Code lines | ≈ 40 | ≈ 45 | ≈ 80 |
| Learned | Pipeline skeleton, basic evaluation | Feature engineering, trustworthy evaluation, diagnostics | Training loop internals, gradient descent in practice |
| Black-box level | fit is a black box | Still black box, but validation makes you trust it | No black box, all handwritten |
| Applicable stage | First baseline | Engineering iteration, pre-deployment | Switching engine, scenarios beyond structured data |
The most profound comparison of the three versions: V1 and V3 use the same random split (random_state=42, same 30% test set), and accuracy goes from 0.933 to 0.978. Where does the improvement come from? Stronger model, standardization, more training epochs — but not from the data (same dataset). This reminds you: the ceiling of modeling is determined by data, and the model's job is to extract as much signal from data as possible. To understand the principles behind every column, read in order: Feature Engineering, Model Evaluation and Validation, Deep Learning Fundamentals.
VI. Debugging and Expansion Directions
1. Most Common Pitfalls and Troubleshooting for All Three Versions
| Symptom | Possible Cause | Troubleshooting |
|---|---|---|
| V1 accuracy below 0.9 | Features not standardized + logistic regression L2 regularization skewed by large-scale features | Add StandardScaler, or increase max_iter |
| V2 train curve 1.0 but val low (trumpet shape) | Overfitting: polynomial dimension too high / trees too deep | Decrease degree, limit tree depth, add regularization |
| V2 both curves low | Underfitting: insufficient model capacity | Switch to stronger model or add features |
| V3 loss not dropping (stuck around 1.1) | Learning rate too large/too small, or not standardized | Adjust lr (0.1→0.01→0.001), check scaler |
| V3 loss drops sharply then train 1.0, test poor | Overfitting: hidden layer too wide / too many epochs | Decrease hidden_dim, add Dropout, early stopping |
| Everything breaks on a new dataset | Missing data cleaning, class imbalance | First check class distribution and missing values |
Universal debugging iron rule: split the problem — first confirm data is fine (print X.shape, y distribution, first few samples), then confirm the pipeline is fine (test on the training set once; if the training set can't be learned either, the problem is the model, not the validation method), and only then tune hyperparameters.
2. Expansion Directions (ranked by cost-effectiveness)
① Hyperparameter tuning (biggest ROI): replace V2's random forest with grid search GridSearchCV, or make V3's lr, hidden_dim, epochs a search space. Full methodology for tuning (search strategies, early stopping, random search) is in Tuning Practice.
② Switch to real data: Iris is too clean. Switch to a real dataset with noise, missing values, and class imbalance, and 80% of the skills you've learned will be exposed at the "data cleaning" stage. Dataset list is in Datasets and Tools Archive.
③ Add visualization: use matplotlib to plot V3's loss curve and V2's learning curves side by side; use PCA to reduce 4D features to 2D and plot scatter plots, visually seeing the distribution of the three classes and the model's decision boundary.
④ From classification to regression: replace Iris with California housing (sklearn.datasets.fetch_california_housing), and rewrite all three versions — the only difference between classification and regression is the loss function (CrossEntropyLoss → MSELoss) and evaluation metric (accuracy → RMSE); the rest of the pipeline is almost identical.
⑤ Complete the engineering: save V2's Pipeline with joblib, write train.py / predict.py, add --help CLI arguments — this is the starting point of Building an ML Project from Scratch, and the first step toward production.
3. What to Do Next
| Your Current State | Recommended Path |
|---|---|
| V1 just ran | Read Model Evaluation and Validation, then come back to V2 |
| V2 ran and can explain learning curves | Read Feature Engineering + Tuning Practice, then switch to real data |
| V3 ran and can modify network structure | Read Deep Learning Fundamentals, try deeper networks, add Dropout, change loss functions |
VII. Further Reading
On-site (shallow to deep)
- What is Machine Learning — the theoretical foundation of this article, read this first to build the big picture
- Supervised Learning — the paradigm all three versions belong to
- Model Evaluation and Validation — the full theory of V2 cross-validation and learning curves
- Feature Engineering — the systematic methods of V2 feature engineering
- Optimization and Gradient Descent — the math behind V3's training loop
- Deep Learning Fundamentals — V3 MLP extensions (CNNs, RNNs, Transformer starting point)
- Tuning Practice — going from V2/V3 to automatic tuning
- Common Pitfalls — the full version of this article's "Debugging and Expansion" section
- Datasets and Tools Archive — your next step after replacing Iris
Off-site references (all real resources)
- scikit-learn official user guide — first-hand authoritative docs for
LogisticRegression,Pipeline,learning_curve,StratifiedKFold - scikit-learn: Iris dataset docs — official documentation and citation info for this article's dataset
- PyTorch official tutorial: Deep Learning with PyTorch: A 60 Minute Blitz — the official companion tutorial for V3, covering tensors, autograd,
nn.Module, and training loops - PyTorch docs: nn.CrossEntropyLoss — official documentation for the loss function used in V3 (including the important detail of built-in softmax)
- Fisher (1936). The use of multiple measurements in taxonomic problems — the original paper for the Iris dataset, the most frequently cited dataset source in ML history
- Andrew Ng. Machine Learning Yearning — free ebook teaching "data, validation set, error analysis" engineering decision-making
- Aurélien Géron. Hands-On Machine Learning with Scikit-Learn, Keras & TensorFlow (O'Reilly) — a practical textbook highly aligned with this article's three-version route, more systematic
Recommended order: run the code for all three versions of this article → read the on-site links you "know what but not why" → open the PyTorch 60-minute tutorial to verify your understanding of V3 → finally pick a real dataset and rewrite all three versions. Between step 3 and step 4, you're already someone who can independently run a complete project end-to-end.