Theme
Progressive Tutorial: Three Iterations to Make It Work
In one sentence: The right way to write deep learning code is not "get it perfect in one shot," but "get it done in three steps" — first a minimal viable version, then add tricks one by one, finally swap the architecture and run full evaluation. Each version introduces only one or two variables, so you always know what each change brought.
This article uses CIFAR-10 (color 32×32, 50,000 training images, 10,000 test images, 10 classes) for the same task, iterating three versions:
| Version | Model | Training Tricks | Test Accuracy (reference) | Training Time (single GPU) |
|---|---|---|---|---|
| v1 | MLP (1 hidden layer only) | None | ~40% | A few minutes |
| v2 | 4-layer CNN | Augmentation + BatchNorm + Adam + Cosine annealing + Early stopping | ~78% | ~20 minutes |
| v3 | Pre-trained ResNet-18 fine-tuned | Transfer learning + Fine-grained learning rates | ~90%+ | ~15 minutes |
This isn't a comparison of "better models," but of smarter engineering decisions. Foundational concepts (backpropagation, loss, optimizers) are covered in Neural Network Fundamentals and Optimization and Gradient Descent; full details on data loading and normalization are in Building a Deep Learning Project from Scratch.
1. v1: Minimal Viable Version
The only goal: verify the "data → model → train → evaluate" pipeline can run. Don't chase performance — go for the shortest path.
python
import torch, torch.nn as nn, torch.nn.functional as F
from torch.utils.data import DataLoader
from torchvision import datasets, transforms
# Data: only the bare minimum ToTensor normalization
transform = transforms.ToTensor()
train_set = datasets.CIFAR10(root="./data", train=True, download=True, transform=transform)
test_set = datasets.CIFAR10(root="./data", train=False, download=True, transform=transform)
train_loader = DataLoader(train_set, batch_size=64, shuffle=True, num_workers=2)
test_loader = DataLoader(test_set, batch_size=256, shuffle=False, num_workers=2)
# Model: single-hidden-layer MLP, ~30 lines
class MLP(nn.Module):
def __init__(self, num_classes=10):
super().__init__()
self.fc1 = nn.Linear(32 * 32 * 3, 512)
self.fc2 = nn.Linear(512, num_classes)
def forward(self, x):
x = x.view(x.size(0), -1)
x = F.relu(self.fc1(x))
return self.fc2(x)
model = MLP()
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9)
for epoch in range(10):
model.train()
total_loss, correct, total = 0.0, 0, 0
for x, y in train_loader:
optimizer.zero_grad()
loss = criterion(model(x), y)
loss.backward()
optimizer.step()
total_loss += loss.item() * x.size(0)
correct += (model(x).argmax(1) == y).sum().item()
total += y.size(0)
print(f"epoch {epoch+1} | loss {total_loss/total:.3f} | acc {correct/total:.3f}")What this version solves: It establishes a baseline, proving that the code, environment, and data flow are all correct. At ~30 lines, any error can be quickly pinpointed.
What the trade-off is: Performance is poor (~40%, not much better than random guessing at 10%). The reason isn't just that the model is too shallow — more critically, it flattens all pixels — CIFAR-10's image structure (colors, edges, textures) is all noise to an MLP. There's also no normalization or augmentation, making training fragile. The full list of pitfalls is in Common Pitfalls and Anti-Patterns.
Why write bad code first
Many people feel embarrassed about having a "bad" baseline. But precisely because it's bad, every subsequent improvement produces a measurable increment. If you start with a pre-trained ResNet and get 90%, you can't tell how much comes from architecture, how much from tricks, and how much from data.
2. v2: Add Training Tricks
With "the pipeline works" as the foundation, this version does one thing: systematically add proven training tricks, running after each addition to see its effect.
python
import torch, torch.nn as nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms
# 1) Data augmentation: random crop + horizontal flip (CIFAR-10 standard config)
train_transform = transforms.Compose([
transforms.RandomCrop(32, padding=4), # pad 4 pixels on all sides, then random-crop back to 32x32
transforms.RandomHorizontalFlip(), # 50% probability horizontal flip
transforms.ToTensor(),
transforms.Normalize((0.4914, 0.4822, 0.4465), (0.2023, 0.1994, 0.2010)),
])
test_transform = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize((0.4914, 0.4822, 0.4465), (0.2023, 0.1994, 0.2010)),
])
train_set = datasets.CIFAR10(root="./data", train=True, download=True, transform=train_transform)
test_set = datasets.CIFAR10(root="./data", train=False, download=True, transform=test_transform)
train_loader = DataLoader(train_set, batch_size=128, shuffle=True, num_workers=2)
test_loader = DataLoader(test_set, batch_size=256, shuffle=False, num_workers=2)
# 2) 4-layer convolution + BatchNorm
class CNN(nn.Module):
def __init__(self, num_classes=10):
super().__init__()
self.features = nn.Sequential(
nn.Conv2d(3, 64, 3, padding=1), nn.BatchNorm2d(64), nn.ReLU(),
nn.Conv2d(64, 64, 3, padding=1), nn.BatchNorm2d(64), nn.ReLU(),
nn.MaxPool2d(2), # 32->16
nn.Conv2d(64, 128, 3, padding=1), nn.BatchNorm2d(128), nn.ReLU(),
nn.Conv2d(128, 128, 3, padding=1), nn.BatchNorm2d(128), nn.ReLU(),
nn.MaxPool2d(2), # 16->8
nn.AdaptiveAvgPool2d(1), # -> [B,128,1,1]
)
self.classifier = nn.Sequential(nn.Flatten(), nn.Linear(128, num_classes))
def forward(self, x):
return self.classifier(self.features(x))
model = CNN()
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
# 3) Learning rate scheduling: cosine annealing (see training recipe for mechanics)
scheduler = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, T_max=20)
# 4) Early stopping: monitor on validation set, stop if no improvement for 5 epochs
best_val_acc, patience, bad_epochs = 0.0, 5, 0
for epoch in range(20):
model.train()
for x, y in train_loader:
optimizer.zero_grad()
loss = criterion(model(x), y)
loss.backward()
optimizer.step()
scheduler.step()
# Validation (for brevity in this teaching example, we use the test set as validation;
# in a real experiment, split off a validation set from the training data,
# following the "touch test set only once" discipline in Evaluation in Practice)
model.eval()
val_correct, val_total = 0, 0
with torch.no_grad():
for x, y in test_loader:
val_correct += (model(x).argmax(1) == y).sum().item()
val_total += y.size(0)
val_acc = val_correct / val_total
print(f"epoch {epoch+1} | val_acc {val_acc:.4f}")
if val_acc > best_val_acc:
best_val_acc = val_acc
bad_epochs = 0
torch.save(model.state_dict(), "best.pt")
else:
bad_epochs += 1
if bad_epochs >= patience:
print("early stop")
breakWhat this version solves:
- Data augmentation directly fights overfitting — each image looks slightly different every epoch, effectively "enlarging" the dataset (principles in Overfitting and Regularization).
- BatchNorm stabilizes the input distribution at each layer, allowing larger learning rates and less worry about initialization (see Initialization and Normalization).
- Adam is insensitive to learning rate choices, converges quickly, and is a safe default when "you don't know what to tune."
- Cosine annealing brings the learning rate very low in later training, helping the loss "settle" near the minimum.
- Early stopping monitors the validation set and automatically decides when training should end, saving time and preventing overfitting.
What the trade-off is:
- Augmentation slows down data loading; if
num_workersisn't sufficient, the GPU will starve. - BatchNorm introduces two new traps: "training/inference behavior mismatch" and "failure when batch size is too small."
- A 4-layer CNN goes from minutes-level to twenty-minutes-level at ~5 min/10 epochs, and Adam's generalization can be worse than well-tuned SGD + momentum on some tasks.
- The validation set is being repeatedly used to "pick" the model — strictly speaking, you should set aside another data split, or at least remember: the accuracy reported here is validation performance, not your final score. Evaluation discipline is detailed in Evaluation in Practice.
v2 test accuracy reaches about 78%. The direction is right, but far from SOTA.
3. v3: Swap Architecture + Full Evaluation
v2's CNN was "trained from scratch," while a ResNet-18 pre-trained on ImageNet has already learned "edges → textures → parts" visual representations (see Representation Learning and Pre-training). Transfer learning puts small dataset tasks right on the shoulders of giants:
python
import torch, torch.nn as nn
from torchvision import models, transforms
from torchvision.datasets import CIFAR10
from torch.utils.data import DataLoader
# Data: use ImageNet mean and std (paired with pre-trained weights, don't use your own statistics)
train_transform = transforms.Compose([
transforms.RandomCrop(32, padding=4),
transforms.RandomHorizontalFlip(),
transforms.ToTensor(),
transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
test_transform = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
train_set = CIFAR10(root="./data", train=True, download=True, transform=train_transform)
test_set = CIFAR10(root="./data", train=False, download=True, transform=test_transform)
train_loader = DataLoader(train_set, batch_size=128, shuffle=True, num_workers=2)
test_loader = DataLoader(test_set, batch_size=256, shuffle=False, num_workers=2)
# Pre-trained ResNet-18: replace the final FC layer with 10 classes
model = models.resnet18(weights=models.ResNet18_Weights.IMAGENET1K_V1)
model.fc = nn.Linear(model.fc.in_features, 10)
# Freeze backbone: Phase 1 — train only the classification head
for param in model.parameters():
param.requires_grad = False
for param in model.fc.parameters():
param.requires_grad = True
optimizer = torch.optim.Adam(model.fc.parameters(), lr=1e-3)
criterion = nn.CrossEntropyLoss()
# Phase 1: train only the classification head (~2-3 epochs is enough)
for epoch in range(3):
model.train()
for x, y in train_loader:
optimizer.zero_grad()
loss = criterion(model(x), y)
loss.backward()
optimizer.step()
print(f"head-only epoch {epoch+1} done")
# Phase 2: unfreeze all parameters, fine-tune with a small learning rate
for param in model.parameters():
param.requires_grad = True
optimizer = torch.optim.Adam(model.parameters(), lr=1e-4) # Fine-tuning lr should be an order of magnitude smaller
for epoch in range(5):
model.train()
for x, y in train_loader:
optimizer.zero_grad()
loss = criterion(model(x), y)
loss.backward()
optimizer.step()
print(f"fine-tune epoch {epoch+1} done")What this version solves:
- Solves "not enough data": CIFAR-10 has only 50K images; training deep networks from scratch easily overfits. Pre-trained features let the model only learn "how to combine features into 10 classes."
- Solves "training takes too long": The freezing phase only updates the 512×10 fully connected layer, which is very fast.
- Transfer learning recipe (freeze → unfreeze, learning rate one order of magnitude apart) — full rules in Training Recipes and Hyperparameter Tuning.
What the trade-off is:
- Resolution mismatch: ResNet was trained on 224×224; feeding 32×32 wastes part of its capability (use
torchvision's resize to 224, or use acifar10-specific pre-trained model instead). - Fine-tuning learning rate can't be large: Pre-trained weights are already great; too large a learning rate will "blow them out" — this is the most common fine-tuning pitfall, see Common Pitfalls and Anti-Patterns.
- Downloading pre-trained weights requires internet, and ResNet-18 is about 45MB.
v3 test accuracy reaches 90%+.
Full Evaluation Report
v3's closing act is "full evaluation" — not just one accuracy number, but a systematic answer to "does this model actually work?":
python
from sklearn.metrics import classification_report, confusion_matrix
import numpy as np
model.eval()
all_pred, all_true = [], []
with torch.no_grad():
for x, y in test_loader:
all_pred.append(model(x).argmax(1).numpy())
all_true.append(y.numpy())
all_pred, all_true = np.concatenate(all_pred), np.concatenate(all_true)
print(classification_report(all_true, all_pred, digits=4))
print(confusion_matrix(all_true, all_pred))In the report you'll see class-imbalanced distributions — classes like "cat" vs "bird" that are hard to distinguish by texture will have lower recall. A real-world deployment model needs to re-weight metrics according to business costs — this is exactly the theme of the entire Evaluation in Practice article.
4. Three-Version Comparison and Philosophy
| Dimension | v1 | v2 | v3 |
|---|---|---|---|
| Model | Single-hidden-layer MLP | 4-layer CNN + BN | Pre-trained ResNet-18 |
| Data | No augmentation | Random crop/flip + normalization | Same as v2 (ImageNet statistics) |
| Optimizer | SGD | Adam + Cosine annealing | Adam (two-phase) |
| Early stopping | No | Yes | Yes |
| Code size | ~30 lines | ~70 lines | ~60 lines + pre-trained |
| Test accuracy | ~40% | ~78% | ~90%+ |
| Main risks | Pipeline doesn't work | New pitfalls from new tricks | Misconfigured transfer, inflated scores |
Taken together, the three versions teach you a philosophy that's more important than any single trick:
- Get the pipeline working first, then make it good — v1's working state is the prerequisite for everything else.
- Introduce one group of variables at a time — every improvement from v2 to v3 can be attributed to specific changes. This is the core of experiment discipline (see DL Design Principles).
- Stop when gains diminish — going from 40% to 78% cost 20 minutes; going from 78% to 90% required transfer learning; every additional percentage point beyond that demands exponential investment.
Further Reading
- Building a Deep Learning Project from Scratch — complete breakdown of v1 and repository structure
- Training Recipes and Hyperparameter Tuning — specific parameters for augmentation, scheduling, and fine-tuning
- Debugging and Diagnostics — how to troubleshoot v2's BatchNorm, early stopping issues
- Evaluation in Practice — how to upgrade v3's evaluation report to a credible assessment
- Representation Learning and Pre-training — theoretical foundation of transfer learning
- CNNs and Computer Vision — evolution of convolutional architectures
References
- Krizhevsky. Learning Multiple Layers of Features from Tiny Images (2009) — Original CIFAR-10 technical report
- He, Zhang, Ren, Sun. Deep Residual Learning for Image Recognition (CVPR 2016) — ResNet paper
- PyTorch. Transfer Learning for Computer Vision Tutorial — Official transfer learning tutorial