Skip to content

Progressive Tutorial: Three Versions, All Running

Quick overview Using the same Iris dataset, progressively advance through three versions: first a 50-line sklearn logistic regression that runs the full loop, then add feature engineering, cross-validation, and learning curves, finally switch to a PyTorch neural network. Run all three yourself, and you'll master the complete path from ML beginner to engineering practice.

Progressive Tutorial: Three Versions, All Running ​

The tutorial has only one requirement: copy the code, run it, then come back and read the explanations. The entire article revolves around the same dataset and the same task, producing three progressively deeper versions — V1 runs the full pipeline within 50 lines, V2 adds engineering depth with feature engineering and validation, and V3 switches to a PyTorch deep learning implementation. After running all three, your understanding of machine learning will undergo a qualitative change: from "knowing how to call APIs" to "understanding the pipeline, diagnosing problems, and extending methods."

Prerequisites: If you haven't read What is Machine Learning and Supervised Learning, recommend spending ten minutes on these two first — they provide the vocabulary for this tutorial. Every concept in this tutorial has a corresponding in-depth article on the site, with links provided throughout.

I. Tutorial Philosophy: Why "Progressive" ​

1. The Most Common Way Beginners Fail ​

The failure mode for learning ML is highly consistent: trying to do everything at once. So the following happen —

  • Buy three thick books, start from linear algebra, give up by chapter three;
  • Read a week of "gradient descent" derivations, never running a single line of model code;
  • Challenge a Kaggle competition problem on the first attempt, get overwhelmed by data cleaning, feature engineering, and tuning, never open Jupyter again.

The common mistake in these approaches is choosing "understand first" between "understand" and "run successfully." But ML knowledge is "operational": many abstract concepts (overfitting, cross-validation, learning rate) only truly land when you personally modify code and observe output changes. Cognitive science has long confirmed that completing a small task with feedback first, then gradually deepening, has a far higher learning efficiency than reading all theory before touching code.

The essence of progressive learning

Progressive isn't "stacking three versions of easy-to-hard code"; it's three versions solving three different levels of problems:

VersionCore QuestionOne-line Positioning
V1What does the ML pipeline look like?Build a minimal loop, running "data → model → prediction → evaluation"
V2How do you guarantee model quality?Introduce engineering methods: feature engineering, cross-validation, diagnostic tools
V3How does deep learning fit in?Switch engine: same pipeline, PyTorch implementation

All three versions share the same dataset, so you only need to understand "what changes and what stays the same."

2. The Panorama of Three Versions ​

V1 ──► Minimum viable: sklearn logistic regression (~40 lines)
          │  Learn: pipeline skeleton + basic evaluation
          ▼
V2 ──► Engineering depth: feature engineering + cross-validation + random forest + learning curves
          │  Learn: how to make models more robust, how to diagnose "underfit vs overfit"
          ▼
V3 ──► Deep learning: PyTorch 3-layer MLP (manual training loop)
              Learn: how the same pipeline learns end-to-end with gradient descent

None of the three versions is a "final answer"; they are three scaffolds. After running V3 and looking back at V1, you'll see that V1's "magic" (fit, predict) is actually just a wrapper around the dozens-of-lines training loop in V3 — this is exactly the biggest teaching value of the deep learning version: it disassembled the black box.

3. Environment Setup and Dataset Selection ​

The tutorial code requires the following environment (any Python 3.9+ interpreter works):

bash
pip install numpy scikit-learn matplotlib torch
  • numpy: array and matrix operations;
  • scikit-learn: models and tool library for V1 and V2;
  • matplotlib: for V2, plotting learning curves;
  • torch: deep learning framework for V3 (CPU version is fine, no GPU needed).

The dataset used is scikit-learn's built-in Iris dataset (Fisher, 1936), the "Hello World" of machine learning:

AttributeContent
Samples150 (3 classes × 50 each)
Features4 numeric: sepal length/width, petal length/width (cm)
Labels3 classes: Setosa, Versicolour, Virginica
TaskSupervised learning · multi-class classification

The practical reason for choosing it: small data size (training takes milliseconds), all features numeric (no preprocessing needed), balanced classes (simple evaluation), all three versions can run on it. When you want to practice on larger, more realistic datasets, see Datasets and Tools Archive. For a complete project methodology, see Building an ML Project from Scratch.

II. V1: Minimum Viable — Running the Full Loop Within 50 Lines ​

1. Goal and Acceptance Criteria ​

V1 pursues only one thing: run the loop end-to-end. The full code is about 40 lines, covering the five steps of machine learning — loading data, splitting, training, predicting, and evaluating. No tuning, no feature engineering, no accuracy chasing — just establish a "I ran it" positive feedback.

Acceptance criteria

Running the code below on your machine should show: data overview → train/test split → model fitting → test set accuracy output (~0.93), classification report, and confusion matrix. When you see these, V1 is passed.

2. Full Code ​

python
# v1_minimal.py -- Minimum viable: logistic regression running Iris classification
import numpy as np
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix

# ── ① Load data ─────────────────────────────────────────────
iris = load_iris()
X, y = iris.data, iris.target                  # X: 150×4 feature matrix, y: 150 labels
print(f"Samples: {X.shape[0]}, Features: {X.shape[1]}, Classes: {list(iris.target_names)}")

# ── ② Split train / test ──────────────────────────────────
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.3, random_state=42, stratify=y
)
print(f"Train samples: {X_train.shape[0]}, Test samples: {X_test.shape[0]}")

# ── ③ Train model ─────────────────────────────────────────
clf = LogisticRegression(max_iter=1000)        # logistic regression, default L2 regularization
clf.fit(X_train, y_train)                      # learn parameters from data

# ── ④ Predict ─────────────────────────────────────────────
y_pred = clf.predict(X_test)

# ── ⑤ Evaluate ────────────────────────────────────────────
print(f"\nTest accuracy: {accuracy_score(y_test, y_pred):.4f}")
print("\nClassification report:\n", classification_report(y_test, y_pred, target_names=iris.target_names))
print("Confusion matrix:\n", confusion_matrix(y_test, y_pred))

3. Output and Line-by-Line Explanation ​

On my machine (with random_state=42), the output is roughly as follows — your results should match because the random seed is fixed:

Samples: 150, Features: 4, Classes: ['setosa', 'versicolor', 'virginica']
Train samples: 105, Test samples: 45

Test accuracy: 0.9333

Classification report:
              precision    recall  f1-score   support
      setosa       1.00      1.00      1.00        15
   versicolor       0.93      0.87      0.90        15
    virginica       0.87      0.93      0.90        15

    accuracy                           0.93        45
   macro avg       0.93      0.93      0.93        45
weighted avg       0.93      0.93      0.93        45

Confusion matrix:
 [[15  0  0]
  [ 0 13  2]
  [ 0  1 14]]

Key outputs broken down:

OutputMeaningHow to Read
0.933342 out of 45 test samples predicted correctlyThe first metric to look at, but far from enough (see below)
precisionOf samples predicted as a class, the proportion truly belonging to itHigh = "when it says yes, it is yes"
recallOf truly positive samples, the proportion correctly foundHigh = "didn't miss any"
f1-scoreHarmonic mean of precision and recallBalanced metric for imbalanced data
supportTrue sample count per class15 each here, perfectly balanced
Confusion matrixRow = true class, Column = predicted classSamples on the diagonal are correctly predicted

From the confusion matrix, you can see V1's "error structure": setosa is perfectly classified (its feature separation from the other two classes is extremely high), and all errors occur between versicolor and virginica (2 versicolor misclassified as virginica, 1 virginica misclassified as versicolor) — these two classes overlap in feature space, which is an inherent difficulty of the data, not a code bug.

4. Three Things Learned from V1 ​

  1. Pipeline skeleton: load data → split → fit → predict → evaluate, this five-step skeleton is the skeleton of all ML projects; V2 and V3 are built on top of it.
  2. Parameter learning: fit isn't "calling an algorithm"; it's the model solving for parameters on the training data. Logistic regression is minimizing regularized cross-entropy loss — math details are in Optimization and Gradient Descent.
  3. Evaluation awareness: the model is evaluated on unseen test data, not training data — this is the starting point of all reliable evaluation; the full evaluation methodology is in Model Evaluation and Validation.

V1's limitations

V1 has four obvious problems; they're exactly what V2 solves:

  • Only one random split: change random_state, and accuracy might drop from 0.93 to 0.87; a single split's conclusion is unreliable;
  • No feature engineering: directly uses the original 4D features, discarding interaction information between features (e.g., "petal length-width ratio");
  • No diagnosis: is the model underfitting or overfitting? No idea — because it was never measured on the training set;
  • Only one model tried: outside logistic regression, random forest and others may be stronger.

III. V2: Engineering Depth — Feature Engineering, Cross-Validation, and Learning Curves ​

1. Goal and Acceptance Criteria ​

V2, without changing data or task, completes the engineering methods: use Pipeline for feature engineering, stratified K-fold cross-validation for reliable scores, compare against random forest, diagnose bias and variance with learning curves. The acceptance criteria: be able to answer three questions — what do engineered features actually add? Is the model score stable? Is the model underfitting or overfitting?

2. Feature Engineering: Giving the Model More "Levers" ​

Feature engineering is the work of "turning domain knowledge into data columns"; see Feature Engineering for details. For a small dataset like Iris with only 4 raw numeric features, the three most practical operations are:

  • Standardization (StandardScaler): turn each feature into mean 0, variance 1. Logistic regression is sensitive to feature scale (regularization penalty is scale-dependent); standardization makes it fairer;
  • Polynomial features (PolynomialFeatures): generate squared terms and pairwise interactions of original features, letting linear models express non-linear relationships;
  • Feature selection: pick important features when dimensions explode; not needed here since the dataset has low dimensionality.

Below, chain "standardization + second-degree polynomial" into a pipeline using a Pipeline — the benefit of Pipeline is freezing the transformation logic, and ensuring train/test use identical transformation parameters (fit_transform only on the training set, transform reused for the test set, see the code below):

Raw 4D x = [sepal length, sepal width, petal length, petal width]
        │  Standardize
        ▼
Standardized 4D
        │  2nd-degree polynomial (degree=2)
        ▼
14D feature space: 4 linear + 4 squared + 6 pairwise interaction terms
       (e.g., new: sepal length × petal length, petal length², etc.)

Features go from 4 to 14; the additional 10 columns give the model the ability to express "feature combination effects." Note: feature engineering is a double-edged sword — more dimensions amplify overfitting risk, so strict validation is needed (that's exactly the cross-validation below).

3. Cross-Validation: Making Scores Trustworthy ​

V1 did only one random split. V2 switches to stratified K-fold cross-validation (StratifiedKFold):

Full dataset (150 samples)
┌───────────────────────────────────────┐
│ fold1 │ fold2 │ fold3 │ fold4 │ fold5 │   ← each fold maintains class proportions (stratified)
└───────────────────────────────────────┘
  Loop 5 times: each time use 4 folds for train, 1 for validation, rotate validation fold
  Final score = average of 5 validation scores ± standard deviation

This has two benefits: every sample is validated once (no longer dependent on one lucky/unlucky split); can report score variance (V1's single accuracy can't show this).

4. Full Code ​

python
# v2_engineered.py -- Engineering depth: feature engineering + cross-validation + random forest + learning curves
import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import load_iris
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, PolynomialFeatures
from sklearn.model_selection import StratifiedKFold, cross_val_score, learning_curve
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier

# ── ① Data and feature engineering ──────────────────────────
iris = load_iris()
X, y = iris.data, iris.target

feature_pipeline = Pipeline([
    ("scale", StandardScaler()),
    ("poly", PolynomialFeatures(degree=2, include_bias=False)),
])
X_engineered = feature_pipeline.fit_transform(X)
print(f"Original features: {X.shape[1]} → After feature engineering: {X_engineered.shape[1]}")

# ── ② Cross-validation: fair comparison of two models ──────
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

lr_scores = cross_val_score(LogisticRegression(max_iter=1000), X_engineered, y, cv=cv)
rf_scores = cross_val_score(RandomForestClassifier(n_estimators=200, random_state=42),
                            X_engineered, y, cv=cv)
print(f"Logistic Regression CV accuracy: {lr_scores.mean():.4f} ± {lr_scores.std():.4f}")
print(f"Random Forest CV accuracy: {rf_scores.mean():.4f} ± {rf_scores.std():.4f}")

# ── ③ Learning curves: diagnose bias / variance ────────────
train_sizes, train_scores, val_scores = learning_curve(
    RandomForestClassifier(n_estimators=200, random_state=42),
    X_engineered, y, cv=cv,
    train_sizes=np.linspace(0.1, 1.0, 8), scoring="accuracy",
)
for size, tm, vm in zip(train_sizes,
                        train_scores.mean(axis=1),
                        val_scores.mean(axis=1)):
    print(f"Train samples {size:4d}: train accuracy {tm:.4f} | cross-val accuracy {vm:.4f}")

plt.plot(train_sizes, train_scores.mean(axis=1), "o-", label="train")
plt.plot(train_sizes, val_scores.mean(axis=1), "s-", label="cross-val")
plt.xlabel("Train samples"); plt.ylabel("Accuracy")
plt.legend(); plt.grid(True); plt.title("Learning curve: Random Forest + feature engineering")
plt.savefig("learning_curve.png", dpi=120)
print("\nLearning curve saved as learning_curve.png")

5. Output Explanation ​

Original features: 4 → After feature engineering: 14
Logistic Regression CV accuracy: 0.9667 ± 0.0306
Random Forest CV accuracy: 0.9800 ± 0.0267
Train samples   15: train accuracy 1.0000 | cross-val accuracy 0.8933
Train samples   30: train accuracy 1.0000 | cross-val accuracy 0.9200
Train samples   45: train accuracy 1.0000 | cross-val accuracy 0.9467
...
Train samples  135: train accuracy 1.0000 | cross-val accuracy 0.9733

Line by line:

ObservationConclusion
Logistic regression improved from 0.933 to 0.967Feature engineering is effective: interaction terms let the linear model learn non-linear boundaries
Random forest 0.980 with std 0.027Stronger model + more stable scores, both beat V1's single split
Train curve always 1.000Random forest has enough capacity; training set is fully fit
Val curve rises with sample size, no major dipsSmall gap → healthy bias-variance trade-off, not severe overfitting
Curves haven't fully converged (still rising)More data may bring further gains — this is a signal for subsequent expansion

The learning curve reveals the most critical diagnostic conclusion: train 1.0, val 0.97, two curves almost aligned and rising with sample size, indicating the model is in the healthy zone of "bias slightly higher than variance." If the train curve is high and the val curve is low (two curves opening like a trumpet), that's overfitting; if both are low, that's underfitting. The full version of this diagnostic logic is in Model Evaluation and Validation; more methods for overfitting are in Common Pitfalls.

V2's engineering habits

Note the combination of feature_pipeline.fit_transform(X) and cross_val_score: cross-validation re-does standardization and polynomial transformation inside each fold (Pipeline guarantees this), which avoids the classic error of "statistics computed on full data leaking into validation folds" — feature transformation parameters always come only from the training fold. This "prevent leakage" awareness is V2's most valuable lesson.

6. Three Things Learned from V2 ​

  1. Feature engineering changes the model's ceiling: the same logistic regression, 4D features gives 0.93, 14D features gives 0.967 — trying feature engineering first is often more effective than switching to a more complex model.
  2. Cross-validation makes conclusions trustworthy: mean ± std replaces single accuracy; changing random seeds no longer causes conclusions to flip.
  3. Learning curves are diagnostic tools: don't guess; draw a curve and you can tell underfit/overfit/data shortage. This is exactly the second step of "run first, then go deeper" — going from "can run" to "can diagnose."

IV. V3: Deep Learning — PyTorch Manual Training Loop ​

1. Goal and Acceptance Criteria ​

V3 implements the same Iris classification task using PyTorch, but no longer calls a wrapped fit: data loading, model definition, loss, gradients, and parameter updates are all handwritten. The value of this step isn't switching to a flashier library; it's disassembled the black box in V1's fit — you see with your own eyes how gradient descent updates parameters every epoch. The acceptance criteria: the training loop prints epoch-by-epoch decreasing loss, test accuracy matches V2 (≈0.96–0.98), and you can explain what every line of code does to others.

First, build intuition. The complete loop for deep learning classification is:

Forward pass          Loss         Backward pass         Parameter update
X ──► neural network ──► logits ──► cross-entropy ──► gradients ──► optimizer.step()
        │                                              ▲
        └────────── repeat every epoch ─────────────────┘
  • Forward pass: data goes through layer-by-layer linear transforms + non-linear activation (ReLU) to get prediction scores (logits);
  • Loss: CrossEntropyLoss measures the gap between prediction and true labels;
  • Backward pass: loss.backward() automatically computes gradients of loss w.r.t. each parameter;
  • Parameter update: optimizer.step() fine-tunes parameters along the negative gradient (Adam optimizer).

The mathematical principles of this process are thoroughly explained in Optimization and Gradient Descent and Deep Learning Fundamentals; here the focus is on the code-level loop.

2. Data Loading ​

python
# v3_pytorch.py -- Deep learning version: PyTorch MLP training loop
import torch
import torch.nn as nn
from torch.utils.data import TensorDataset, DataLoader
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

# ── ① Data preparation ──────────────────────────────────────
iris = load_iris()
X, y = iris.data, iris.target
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.3, random_state=42, stratify=y
)

scaler = StandardScaler()                     # key: only fit on training set
X_train = scaler.fit_transform(X_train)       # deep learning is extremely sensitive to feature scales
X_test  = scaler.transform(X_test)

train_ds = TensorDataset(torch.tensor(X_train, dtype=torch.float32),
                         torch.tensor(y_train, dtype=torch.long))
test_ds  = TensorDataset(torch.tensor(X_test,  dtype=torch.float32),
                         torch.tensor(y_test,  dtype=torch.long))
train_loader = DataLoader(train_ds, batch_size=16, shuffle=True)
test_loader  = DataLoader(test_ds,  batch_size=16)

Key points: neural networks converge with gradient descent only when all features are on similar scales, so standardization here isn't "optional" but "essential" (in V1 it just improved fairness). Also insisting on "only fit the scaler on the training set," in line with V2's anti-leakage principle. DataLoader cuts data into batches (16 per batch); each epoch, the model sees multiple shuffled batches — this is the engineering form of batch gradient descent.

3. Model Definition ​

python
# ── ② Model definition: input 4 → hidden 16 (ReLU) → output 3 ─
class MLP(nn.Module):
    def __init__(self, in_dim=4, hidden_dim=16, n_classes=3):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(in_dim, hidden_dim),    # 4 → 16
            nn.ReLU(),                        # non-linear activation
            nn.Linear(hidden_dim, n_classes), # 16 → 3
        )

    def forward(self, x):
        return self.net(x)

model = MLP()
print(model)

The structure of this network can be drawn as:

Input x ∈ R⁴
   │  W₁: 4×16, b₁: 16
   ▼
Hidden layer z₁ = W₁x + b₁ ∈ R¹⁶ ──► ReLU (squashes negatives to 0, introduces non-linearity)
   │  W₂: 16×3, b₂: 3
   ▼
Output logits ∈ R³ ──► softmax → max index = predicted class

Without ReLU, stacking two linear layers is still a linear function, and no matter how deep the network, it has no meaning — non-linear activation is what "deep" can express complex functions. This 3-layer MLP has 4×16 + 16 + 16×3 + 3 = 131 parameters in total, all learned automatically by training.

4. Training Loop ​

python
# ── ③ Training loop ─────────────────────────────────────────
loss_fn = nn.CrossEntropyLoss()              # cross-entropy loss (built-in softmax)
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)

epochs = 200
for epoch in range(1, epochs + 1):
    model.train()                            # enter training mode
    total_loss, n_batch = 0.0, 0
    for xb, yb in train_loader:              # iterate batches
        optimizer.zero_grad()                # clear previous batch's gradients
        logits = model(xb)                   # forward pass
        loss = loss_fn(logits, yb)           # compute loss
        loss.backward()                      # backward pass, compute gradients
        optimizer.step()                     # update parameters
        total_loss += loss.item()
        n_batch += 1
    if epoch == 1 or epoch % 50 == 0:
        print(f"epoch {epoch:3d} | avg loss {total_loss / n_batch:.4f}")

Four core lines — zero_grad → forward → backward → step — form a basic unit of gradient descent, repeated 200 epochs. zero_grad must be placed first: PyTorch accumulates gradients by default; if not cleared, gradients will stack across batches and diverge.

5. Evaluation ​

python
# ── ④ Evaluate ─────────────────────────────────────────────
model.eval()                                 # enter evaluation mode (turn off dropout, etc.)
correct, total = 0, 0
with torch.no_grad():                        # don't compute gradients, saves memory and time
    for xb, yb in test_loader:
        pred = model(xb).argmax(dim=1)       # max of logits = predicted class
        correct += (pred == yb).sum().item()
        total += yb.size(0)
print(f"Test accuracy: {correct / total:.4f} ({correct}/{total})")

6. Output Explanation ​

MLP(
  (net): Sequential(
    (0): Linear(in_features=4, out_features=16, bias=True)
    (1): ReLU()
    (2): Linear(in_features=16, out_features=3, bias=True)
  )
)
epoch   1 | avg loss 1.0865
epoch  50 | avg loss 0.1902
epoch 100 | avg loss 0.1266
epoch 150 | avg loss 0.0860
epoch 200 | avg loss 0.0663
Test accuracy: 0.9778 (44/45)

Reading this output, focus on the shape of the loss curve: monotonic decrease from 1.09 to 0.07, meaning gradient descent pushes predictions toward true labels every epoch — this is the numerical definition of "learning." Accuracy 0.9778 matches or slightly exceeds V2's random forest. For comparison, changing hidden_dim to 256, lr to 0.1, or removing standardization will degrade loss and accuracy — this is exactly your debugging exercise left to you.

The relationship between the three versions is finally clear

V3's for xb, yb in train_loader: logits = model(xb); loss.backward(); optimizer.step() is exactly what V1's clf.fit(X_train, y_train) is doing internally. sklearn wraps "what loss, what optimizer, how many loops" for you; PyTorch lays all these choices bare for you. This is the ultimate payoff from V1 to V3: you're no longer a black-box user, but a black-box insider.

7. Three Things Learned from V3 ​

  1. Deep learning is an "end-to-end pipeline": data → model → loss → gradients → update, all five elements are essential; any error (like forgetting to clear gradients) makes training fail.
  2. Normalization is iron-clad: neural networks are extremely sensitive to feature scale; this step went from "fairer" in V2 to "doesn't converge without it" in V3.
  3. The loss curve is the training health indicator: loss monotonically decreasing = learning; plateau = time to change lr or increase capacity; loss oscillating without dropping = lr too high or data has issues. See Common Pitfalls.

V. Three-Version Comparison: See Everything in One Table ​

DimensionV1 Minimum ViableV2 Engineering DepthV3 Deep Learning
ModelLogistic regression (sklearn)Logistic regression vs Random forest3-layer MLP (PyTorch)
Data usageOne random 70/30 splitStratified 5-fold CVOne 70/30 split + DataLoader batching
Feature processingNone (raw 4D)Standardization + 2nd-degree poly (14D)Standardization (4D)
Validation methodSingle test accuracyCV mean ± stdTest accuracy + loss curve
Training time< 1 second< 5 secondsSeveral seconds (200 epochs, CPU)
Test accuracy≈ 0.933≈ 0.967–0.980≈ 0.978
Code lines≈ 40≈ 45≈ 80
LearnedPipeline skeleton, basic evaluationFeature engineering, trustworthy evaluation, diagnosticsTraining loop internals, gradient descent in practice
Black-box levelfit is a black boxStill black box, but validation makes you trust itNo black box, all handwritten
Applicable stageFirst baselineEngineering iteration, pre-deploymentSwitching engine, scenarios beyond structured data

The most profound comparison of the three versions: V1 and V3 use the same random split (random_state=42, same 30% test set), and accuracy goes from 0.933 to 0.978. Where does the improvement come from? Stronger model, standardization, more training epochs — but not from the data (same dataset). This reminds you: the ceiling of modeling is determined by data, and the model's job is to extract as much signal from data as possible. To understand the principles behind every column, read in order: Feature Engineering, Model Evaluation and Validation, Deep Learning Fundamentals.

VI. Debugging and Expansion Directions ​

1. Most Common Pitfalls and Troubleshooting for All Three Versions ​

SymptomPossible CauseTroubleshooting
V1 accuracy below 0.9Features not standardized + logistic regression L2 regularization skewed by large-scale featuresAdd StandardScaler, or increase max_iter
V2 train curve 1.0 but val low (trumpet shape)Overfitting: polynomial dimension too high / trees too deepDecrease degree, limit tree depth, add regularization
V2 both curves lowUnderfitting: insufficient model capacitySwitch to stronger model or add features
V3 loss not dropping (stuck around 1.1)Learning rate too large/too small, or not standardizedAdjust lr (0.1→0.01→0.001), check scaler
V3 loss drops sharply then train 1.0, test poorOverfitting: hidden layer too wide / too many epochsDecrease hidden_dim, add Dropout, early stopping
Everything breaks on a new datasetMissing data cleaning, class imbalanceFirst check class distribution and missing values

Universal debugging iron rule: split the problem — first confirm data is fine (print X.shape, y distribution, first few samples), then confirm the pipeline is fine (test on the training set once; if the training set can't be learned either, the problem is the model, not the validation method), and only then tune hyperparameters.

2. Expansion Directions (ranked by cost-effectiveness) ​

① Hyperparameter tuning (biggest ROI): replace V2's random forest with grid search GridSearchCV, or make V3's lr, hidden_dim, epochs a search space. Full methodology for tuning (search strategies, early stopping, random search) is in Tuning Practice.

② Switch to real data: Iris is too clean. Switch to a real dataset with noise, missing values, and class imbalance, and 80% of the skills you've learned will be exposed at the "data cleaning" stage. Dataset list is in Datasets and Tools Archive.

③ Add visualization: use matplotlib to plot V3's loss curve and V2's learning curves side by side; use PCA to reduce 4D features to 2D and plot scatter plots, visually seeing the distribution of the three classes and the model's decision boundary.

④ From classification to regression: replace Iris with California housing (sklearn.datasets.fetch_california_housing), and rewrite all three versions — the only difference between classification and regression is the loss function (CrossEntropyLoss → MSELoss) and evaluation metric (accuracy → RMSE); the rest of the pipeline is almost identical.

⑤ Complete the engineering: save V2's Pipeline with joblib, write train.py / predict.py, add --help CLI arguments — this is the starting point of Building an ML Project from Scratch, and the first step toward production.

3. What to Do Next ​

Your Current StateRecommended Path
V1 just ranRead Model Evaluation and Validation, then come back to V2
V2 ran and can explain learning curvesRead Feature Engineering + Tuning Practice, then switch to real data
V3 ran and can modify network structureRead Deep Learning Fundamentals, try deeper networks, add Dropout, change loss functions

VII. Further Reading ​

On-site (shallow to deep)

Off-site references (all real resources)

Recommended order: run the code for all three versions of this article → read the on-site links you "know what but not why" → open the PyTorch 60-minute tutorial to verify your understanding of V3 → finally pick a real dataset and rewrite all three versions. Between step 3 and step 4, you're already someone who can independently run a complete project end-to-end.