Skip to content

Progressive Tutorial: Three Iterations to Make It Work

Quick overview The same image classification task, iterated through three code versions: v1 is the minimal runnable MLP to get the pipeline going, v2 adds data augmentation, BatchNorm, learning rate scheduling, and early stopping, v3 switches to a pre-trained CNN and performs full evaluation. Each version clarifies "what problem it solves, what the trade-off is," with a comparison table.

Progressive Tutorial: Three Iterations to Make It Work ​

In one sentence: The right way to write deep learning code is not "get it perfect in one shot," but "get it done in three steps" — first a minimal viable version, then add tricks one by one, finally swap the architecture and run full evaluation. Each version introduces only one or two variables, so you always know what each change brought.

This article uses CIFAR-10 (color 32×32, 50,000 training images, 10,000 test images, 10 classes) for the same task, iterating three versions:

VersionModelTraining TricksTest Accuracy (reference)Training Time (single GPU)
v1MLP (1 hidden layer only)None~40%A few minutes
v24-layer CNNAugmentation + BatchNorm + Adam + Cosine annealing + Early stopping~78%~20 minutes
v3Pre-trained ResNet-18 fine-tunedTransfer learning + Fine-grained learning rates~90%+~15 minutes

This isn't a comparison of "better models," but of smarter engineering decisions. Foundational concepts (backpropagation, loss, optimizers) are covered in Neural Network Fundamentals and Optimization and Gradient Descent; full details on data loading and normalization are in Building a Deep Learning Project from Scratch.

1. v1: Minimal Viable Version ​

The only goal: verify the "data → model → train → evaluate" pipeline can run. Don't chase performance — go for the shortest path.

python
import torch, torch.nn as nn, torch.nn.functional as F
from torch.utils.data import DataLoader
from torchvision import datasets, transforms

# Data: only the bare minimum ToTensor normalization
transform = transforms.ToTensor()
train_set = datasets.CIFAR10(root="./data", train=True, download=True, transform=transform)
test_set  = datasets.CIFAR10(root="./data", train=False, download=True, transform=transform)
train_loader = DataLoader(train_set, batch_size=64, shuffle=True, num_workers=2)
test_loader  = DataLoader(test_set, batch_size=256, shuffle=False, num_workers=2)

# Model: single-hidden-layer MLP, ~30 lines
class MLP(nn.Module):
    def __init__(self, num_classes=10):
        super().__init__()
        self.fc1 = nn.Linear(32 * 32 * 3, 512)
        self.fc2 = nn.Linear(512, num_classes)

    def forward(self, x):
        x = x.view(x.size(0), -1)
        x = F.relu(self.fc1(x))
        return self.fc2(x)

model = MLP()
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9)

for epoch in range(10):
    model.train()
    total_loss, correct, total = 0.0, 0, 0
    for x, y in train_loader:
        optimizer.zero_grad()
        loss = criterion(model(x), y)
        loss.backward()
        optimizer.step()
        total_loss += loss.item() * x.size(0)
        correct += (model(x).argmax(1) == y).sum().item()
        total += y.size(0)
    print(f"epoch {epoch+1} | loss {total_loss/total:.3f} | acc {correct/total:.3f}")

What this version solves: It establishes a baseline, proving that the code, environment, and data flow are all correct. At ~30 lines, any error can be quickly pinpointed.

What the trade-off is: Performance is poor (~40%, not much better than random guessing at 10%). The reason isn't just that the model is too shallow — more critically, it flattens all pixels — CIFAR-10's image structure (colors, edges, textures) is all noise to an MLP. There's also no normalization or augmentation, making training fragile. The full list of pitfalls is in Common Pitfalls and Anti-Patterns.

Why write bad code first

Many people feel embarrassed about having a "bad" baseline. But precisely because it's bad, every subsequent improvement produces a measurable increment. If you start with a pre-trained ResNet and get 90%, you can't tell how much comes from architecture, how much from tricks, and how much from data.

2. v2: Add Training Tricks ​

With "the pipeline works" as the foundation, this version does one thing: systematically add proven training tricks, running after each addition to see its effect.

python
import torch, torch.nn as nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms

# 1) Data augmentation: random crop + horizontal flip (CIFAR-10 standard config)
train_transform = transforms.Compose([
    transforms.RandomCrop(32, padding=4),   # pad 4 pixels on all sides, then random-crop back to 32x32
    transforms.RandomHorizontalFlip(),       # 50% probability horizontal flip
    transforms.ToTensor(),
    transforms.Normalize((0.4914, 0.4822, 0.4465), (0.2023, 0.1994, 0.2010)),
])
test_transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.4914, 0.4822, 0.4465), (0.2023, 0.1994, 0.2010)),
])
train_set = datasets.CIFAR10(root="./data", train=True, download=True, transform=train_transform)
test_set  = datasets.CIFAR10(root="./data", train=False, download=True, transform=test_transform)
train_loader = DataLoader(train_set, batch_size=128, shuffle=True, num_workers=2)
test_loader  = DataLoader(test_set,  batch_size=256, shuffle=False, num_workers=2)

# 2) 4-layer convolution + BatchNorm
class CNN(nn.Module):
    def __init__(self, num_classes=10):
        super().__init__()
        self.features = nn.Sequential(
            nn.Conv2d(3, 64, 3, padding=1), nn.BatchNorm2d(64), nn.ReLU(),
            nn.Conv2d(64, 64, 3, padding=1), nn.BatchNorm2d(64), nn.ReLU(),
            nn.MaxPool2d(2),                        # 32->16
            nn.Conv2d(64, 128, 3, padding=1), nn.BatchNorm2d(128), nn.ReLU(),
            nn.Conv2d(128, 128, 3, padding=1), nn.BatchNorm2d(128), nn.ReLU(),
            nn.MaxPool2d(2),                        # 16->8
            nn.AdaptiveAvgPool2d(1),                # -> [B,128,1,1]
        )
        self.classifier = nn.Sequential(nn.Flatten(), nn.Linear(128, num_classes))

    def forward(self, x):
        return self.classifier(self.features(x))

model = CNN()
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

# 3) Learning rate scheduling: cosine annealing (see training recipe for mechanics)
scheduler = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, T_max=20)

# 4) Early stopping: monitor on validation set, stop if no improvement for 5 epochs
best_val_acc, patience, bad_epochs = 0.0, 5, 0
for epoch in range(20):
    model.train()
    for x, y in train_loader:
        optimizer.zero_grad()
        loss = criterion(model(x), y)
        loss.backward()
        optimizer.step()
    scheduler.step()

    # Validation (for brevity in this teaching example, we use the test set as validation;
    # in a real experiment, split off a validation set from the training data,
    # following the "touch test set only once" discipline in Evaluation in Practice)
    model.eval()
    val_correct, val_total = 0, 0
    with torch.no_grad():
        for x, y in test_loader:
            val_correct += (model(x).argmax(1) == y).sum().item()
            val_total += y.size(0)
    val_acc = val_correct / val_total
    print(f"epoch {epoch+1} | val_acc {val_acc:.4f}")

    if val_acc > best_val_acc:
        best_val_acc = val_acc
        bad_epochs = 0
        torch.save(model.state_dict(), "best.pt")
    else:
        bad_epochs += 1
        if bad_epochs >= patience:
            print("early stop")
            break

What this version solves:

  • Data augmentation directly fights overfitting — each image looks slightly different every epoch, effectively "enlarging" the dataset (principles in Overfitting and Regularization).
  • BatchNorm stabilizes the input distribution at each layer, allowing larger learning rates and less worry about initialization (see Initialization and Normalization).
  • Adam is insensitive to learning rate choices, converges quickly, and is a safe default when "you don't know what to tune."
  • Cosine annealing brings the learning rate very low in later training, helping the loss "settle" near the minimum.
  • Early stopping monitors the validation set and automatically decides when training should end, saving time and preventing overfitting.

What the trade-off is:

  • Augmentation slows down data loading; if num_workers isn't sufficient, the GPU will starve.
  • BatchNorm introduces two new traps: "training/inference behavior mismatch" and "failure when batch size is too small."
  • A 4-layer CNN goes from minutes-level to twenty-minutes-level at ~5 min/10 epochs, and Adam's generalization can be worse than well-tuned SGD + momentum on some tasks.
  • The validation set is being repeatedly used to "pick" the model — strictly speaking, you should set aside another data split, or at least remember: the accuracy reported here is validation performance, not your final score. Evaluation discipline is detailed in Evaluation in Practice.

v2 test accuracy reaches about 78%. The direction is right, but far from SOTA.

3. v3: Swap Architecture + Full Evaluation ​

v2's CNN was "trained from scratch," while a ResNet-18 pre-trained on ImageNet has already learned "edges → textures → parts" visual representations (see Representation Learning and Pre-training). Transfer learning puts small dataset tasks right on the shoulders of giants:

python
import torch, torch.nn as nn
from torchvision import models, transforms
from torchvision.datasets import CIFAR10
from torch.utils.data import DataLoader

# Data: use ImageNet mean and std (paired with pre-trained weights, don't use your own statistics)
train_transform = transforms.Compose([
    transforms.RandomCrop(32, padding=4),
    transforms.RandomHorizontalFlip(),
    transforms.ToTensor(),
    transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
test_transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
train_set = CIFAR10(root="./data", train=True, download=True, transform=train_transform)
test_set  = CIFAR10(root="./data", train=False, download=True, transform=test_transform)
train_loader = DataLoader(train_set, batch_size=128, shuffle=True, num_workers=2)
test_loader  = DataLoader(test_set, batch_size=256, shuffle=False, num_workers=2)

# Pre-trained ResNet-18: replace the final FC layer with 10 classes
model = models.resnet18(weights=models.ResNet18_Weights.IMAGENET1K_V1)
model.fc = nn.Linear(model.fc.in_features, 10)

# Freeze backbone: Phase 1 — train only the classification head
for param in model.parameters():
    param.requires_grad = False
for param in model.fc.parameters():
    param.requires_grad = True

optimizer = torch.optim.Adam(model.fc.parameters(), lr=1e-3)
criterion = nn.CrossEntropyLoss()

# Phase 1: train only the classification head (~2-3 epochs is enough)
for epoch in range(3):
    model.train()
    for x, y in train_loader:
        optimizer.zero_grad()
        loss = criterion(model(x), y)
        loss.backward()
        optimizer.step()
    print(f"head-only epoch {epoch+1} done")

# Phase 2: unfreeze all parameters, fine-tune with a small learning rate
for param in model.parameters():
    param.requires_grad = True
optimizer = torch.optim.Adam(model.parameters(), lr=1e-4)  # Fine-tuning lr should be an order of magnitude smaller

for epoch in range(5):
    model.train()
    for x, y in train_loader:
        optimizer.zero_grad()
        loss = criterion(model(x), y)
        loss.backward()
        optimizer.step()
    print(f"fine-tune epoch {epoch+1} done")

What this version solves:

  • Solves "not enough data": CIFAR-10 has only 50K images; training deep networks from scratch easily overfits. Pre-trained features let the model only learn "how to combine features into 10 classes."
  • Solves "training takes too long": The freezing phase only updates the 512×10 fully connected layer, which is very fast.
  • Transfer learning recipe (freeze → unfreeze, learning rate one order of magnitude apart) — full rules in Training Recipes and Hyperparameter Tuning.

What the trade-off is:

  • Resolution mismatch: ResNet was trained on 224×224; feeding 32×32 wastes part of its capability (use torchvision's resize to 224, or use a cifar10-specific pre-trained model instead).
  • Fine-tuning learning rate can't be large: Pre-trained weights are already great; too large a learning rate will "blow them out" — this is the most common fine-tuning pitfall, see Common Pitfalls and Anti-Patterns.
  • Downloading pre-trained weights requires internet, and ResNet-18 is about 45MB.

v3 test accuracy reaches 90%+.

Full Evaluation Report ​

v3's closing act is "full evaluation" — not just one accuracy number, but a systematic answer to "does this model actually work?":

python
from sklearn.metrics import classification_report, confusion_matrix
import numpy as np

model.eval()
all_pred, all_true = [], []
with torch.no_grad():
    for x, y in test_loader:
        all_pred.append(model(x).argmax(1).numpy())
        all_true.append(y.numpy())
all_pred, all_true = np.concatenate(all_pred), np.concatenate(all_true)

print(classification_report(all_true, all_pred, digits=4))
print(confusion_matrix(all_true, all_pred))

In the report you'll see class-imbalanced distributions — classes like "cat" vs "bird" that are hard to distinguish by texture will have lower recall. A real-world deployment model needs to re-weight metrics according to business costs — this is exactly the theme of the entire Evaluation in Practice article.

4. Three-Version Comparison and Philosophy ​

Dimensionv1v2v3
ModelSingle-hidden-layer MLP4-layer CNN + BNPre-trained ResNet-18
DataNo augmentationRandom crop/flip + normalizationSame as v2 (ImageNet statistics)
OptimizerSGDAdam + Cosine annealingAdam (two-phase)
Early stoppingNoYesYes
Code size~30 lines~70 lines~60 lines + pre-trained
Test accuracy~40%~78%~90%+
Main risksPipeline doesn't workNew pitfalls from new tricksMisconfigured transfer, inflated scores

Taken together, the three versions teach you a philosophy that's more important than any single trick:

  1. Get the pipeline working first, then make it good — v1's working state is the prerequisite for everything else.
  2. Introduce one group of variables at a time — every improvement from v2 to v3 can be attributed to specific changes. This is the core of experiment discipline (see DL Design Principles).
  3. Stop when gains diminish — going from 40% to 78% cost 20 minutes; going from 78% to 90% required transfer learning; every additional percentage point beyond that demands exponential investment.

Further Reading ​

References ​