Theme
CNN and Computer Vision
Concept Definition: Seeing the World with a Locality Prior
Convolutional Neural Networks (CNNs) are neural networks designed specifically for grid-structured data like images. Their core assumption is locality: the relationship between adjacent pixels in an image is far more important than that between distant pixels — a cat's ears, eyes, and whiskers are local features that, through sliding windows (convolutional kernels) extraction, are then composed layer by layer into the semantics of "cat."
Raw image → Convolutional layers (extract local features) → Pooling layers (reduce dimensions)
→ Convolutional layers → Pooling layers → … → flatten → fully connected layers → class probabilitiesTwo decisive advantages of CNNs over fully connected networks:
- Parameter sharing: one convolutional kernel slides across the entire image and reuses — for a 224×224×3 image, the first fully connected layer has tens of millions of parameters, while CNN uses only thousands;
- Translation invariance: whether a cat is in the top-left or bottom-right of the image, the convolutional kernel recognizes it — fully connected networks can't do this.
2. The Three Core Components of CNNs
1. Convolutional layers: sliding feature detectors
A convolutional kernel (e.g., 3×3×3) slides over the input, computing weighted sums at each position, producing a feature map. Multiple kernels = multiple feature maps (detecting different patterns: edges, colors, textures). Two hyperparameters:
- Stride: how many pixels the kernel moves each slide; when >1, the output size shrinks;
- Padding: zero-padding at edges to maintain output size (same padding).
Receptive field: the size of the input region that each output pixel "sees." Shallow layers have small receptive fields (looking locally), deep layers have large receptive fields (looking globally) — this is the mechanism behind "shallow layers learn edges, deep layers learn semantics."
2. Pooling layers: compression and invariance
Max pooling: take the maximum value in a 2×2 window — preserves the strongest response, discards exact position. Benefits: dimensionality reduction (reduces computation), introduces some translation invariance, and prevents overfitting. Modern networks (post-ResNet) tend to use "convolution with stride 2" instead of pooling, but the idea is the same.
3. Residual connections: the key to making networks deeper
Before ResNet in 2015, the dilemma was: as networks grew to 20+ layers, training error actually increased (degradation problem — not overfitting, but deep networks being hard to optimize). ResNet's solution is extremely simple:
output = F(x) + x (F is a convolutional block, x is the input, i.e., the "residual")Let the network learn the "residual" rather than the complete mapping: if a layer is already optimal, residual learning only needs to push F to 0 (output = input), sharply reducing learning difficulty. With this, ResNet reached 152 layers, ImageNet top-5 error of 3.57% — beating humans for the first time (~5.1%). Residual connections are now standard in all deep networks (they also appear in Transformers).
3. Architecture Evolution Timeline
| Year | Model | Key contribution | ImageNet top-5 error |
|---|---|---|---|
| 2012 | AlexNet | GPU + ReLU + Dropout + data augmentation, ignited deep learning | 15.3% (previous year's champion: 26.2%) |
| 2014 | VGG | Stack small kernels (3×3) deeper, clean structure | 7.3% |
| 2014 | GoogLeNet (Inception) | Multi-scale conv parallel + 1×1 conv for dimensionality reduction | 6.7% |
| 2015 | ResNet | Residual connections, 152 layers | 3.57% (beat humans) |
| 2017 | DenseNet | Dense inter-layer connections, feature reuse | 3.46% |
| 2019 | EfficientNet | Neural architecture search (NAS) jointly scaling depth/width/resolution | Below 2.9% |
| 2020 | ViT | Transformer eats image patches directly, start of visual foundation models | Catching up to CNN |
Trend: from "hand-designed architectures" (VGG/ResNet) to "NAS auto-search architectures" (EfficientNet), to "Transformer unifying vision and language" (ViT). Vision foundation models post-2021 (CLIP, SAM, DINO) are mostly based on ViT or CNN hybrids, but CNN's locality prior still dominates lightweight, mobile scenarios.
4. Image Task Families
| Task | Goal | Representative models | Metrics |
|---|---|---|---|
| Image classification | Whole image → class | ResNet, EfficientNet, ViT | top-1/top-5 accuracy |
| Object detection | Locate + classify multiple targets | YOLO (one-stage), Faster R-CNN (two-stage) | mAP |
| Semantic segmentation | Each pixel → class | U-Net, DeepLab | mIoU |
| Instance segmentation | Segment + distinguish individuals | Mask R-CNN | mAP |
| Pose estimation | Key point localization | HRNet, OpenPose | OKS/AP |
| Image retrieval / embedding | Image → vector | Metric learning, CLIP | Recall@K |
Detection vs. segmentation: detection gives "box + class"; segmentation gives "pixel-level mask." Business use: classification/detection for quality inspection, segmentation for medical imaging and autonomous driving, embedding for cross-modal search.
5. The Golden Recipe for Image Tasks: Transfer Learning
Almost all practical image projects start from pre-trained models, not training from scratch:
- Load pre-trained weights: ResNet/EfficientNet trained on ImageNet (free, high quality);
- Replace the classification head: swap the final fully connected layer to your task's output dimension;
- Two-stage training: first freeze the backbone and train only the new head (a few epochs), then unfreeze the backbone with a small learning rate for fine-tuning (few epochs);
- Data augmentation: random cropping, flipping, color jittering (feeding data with different variations each epoch, equivalent to free data expansion).
python
import torchvision.models as models
import torch.nn as nn
backbone = models.resnet18(weights=models.ResNet18_Weights.IMAGENET1K_V1)
backbone.fc = nn.Linear(512, num_classes) # Replace classification head
# Freeze backbone (train only the classification head first)
for p in backbone.parameters():
p.requires_grad = False
for p in backbone.fc.parameters():
p.requires_grad = TrueWhy it works: low-level features (edges, textures, shapes) learned on ImageNet are universal; small dataset tasks only need to adapt high-level semantics. The freeze + fine-tune layered strategy is the core playbook for small-data image projects.
Three disciplines of transfer learning
- Don't overdo data augmentation: flipping/cropping destroys label semantics for "direction-sensitive" tasks (license plate recognition, medical images);
- Fine-tuning learning rate must be small: pre-trained weights are already good; a large learning rate ruins them in one step;
- Class imbalance must still be handled: pre-training solves "features," not "sample distribution."
6. Key Techniques for Training Image Models
- BatchNorm: standard practice, stabilizes training, accelerates convergence, with built-in mild regularization (placed after conv, before activation);
- Learning rate: warmup + cosine annealing, peak ~1e-3 (for fine-tuning: 1e-4~1e-5);
- Label smoothing: soften classification head output (don't push to 100% confidence), suppresses overfitting;
- EMA (exponential moving average): average the weights over time, use the averaged weights for inference, stable point gains;
- Mixed precision training: FP16 forward/backward + FP32 master weights, halve VRAM, double speed (PyTorch AMP, one line to enable).
7. Visual Foundation Models and Current State (2023–2026)
After CNNs, vision enters the "foundation model" era:
- CLIP (2021): jointly train image + text; the image encoder learns cross-modal semantics — "photo of a cat" and "cat" map to similar vectors, powering text-to-image (Stable Diffusion uses it for text encoding) and zero-shot classification;
- SAM (2023): Segment Anything, prompt-driven (points/boxes/text) universal segmentation model;
- Vision-Language Models (VLM): GPT-4V, Gemini, Qwen-VL, etc., fold vision understanding into LLMs; image QA and document understanding become mainstream;
- Visual tokenizer trend: more systems cut images into tokens to feed to Transformers — CNN's "local prior" is being replaced by "data scale."
Current state assessment for practitioners: industry vision projects = pre-trained CNN/ViT fine-tuning + detection/segmentation frameworks (MMDetection, Detectron2) + data augmentation + distillation (large models distilled to small models for deployment). Foundation models (VLMs) handle "understanding," small models handle "high-speed inference."
8. Tradeoffs and Decision Points
- CNN vs. ViT: small data, limited resources → CNN (strong prior, efficient); large data, cross-modal → ViT/visual foundation models;
- Accuracy vs. speed: YOLO (fast) or Faster R-CNN (accurate) for detection — decide based on real-time requirements; deploy quantization + TensorRT/ONNX acceleration;
- Fine-tuning vs. distillation: directly fine-tuning a small model is simple; distilling a large model to a small model has higher quality but heavier engineering;
- Supervised vs. self-supervised: when labeling is expensive, use self-supervised pre-training (MAE, DINO) or CLIP zero-shot, then fine-tune with a small amount of labeled data.
Further Reading
- Deep Learning Foundations — The math and network foundation of CNNs
- Transformer and NLP — The architectural source of ViT
- Large Language Models (LLM) — The fusion of vision-language models
- Generative Models — Diffusion models and text-to-image
- Classic Paper Deep Dives — AlexNet/ResNet deep reads
- Datasets and Tools — Public datasets like ImageNet, COCO
References
- Krizhevsky et al. ImageNet Classification with Deep Convolutional Neural Networks (AlexNet, NeurIPS 2012)
- He et al. Deep Residual Learning for Image Recognition (ResNet, CVPR 2016)
- Simonyan & Zisserman. Very Deep Convolutional Networks (VGG, ICLR 2015)
- Tan & Le. EfficientNet: Rethinking Model Scaling (ICML 2019)
- Dosovitskiy et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT, ICLR 2021)
- Redmon et al. You Only Look Once (YOLO, CVPR 2016)
- Radford et al. Learning Transferable Visual Models From Natural Language Supervision (CLIP, ICML 2021)
- Kirillov et al. Segment Anything (SAM, ICCV 2023)
- Deng et al. ImageNet: A Large-Scale Hierarchical Image Database (CVPR 2009) — ImageNet dataset paper