Skip to content

CNN and Computer Vision

Quick overview From AlexNet to ViT — how convolutional neural networks dominated computer vision for a decade. This article dissects CNN architecture evolution (AlexNet/ResNet/EfficientNet), image task families (classification/detection/segmentation), training tricks (transfer learning/data augmentation), and the current state of visual foundation models.

CNN and Computer Vision ​

Concept Definition: Seeing the World with a Locality Prior ​

Convolutional Neural Networks (CNNs) are neural networks designed specifically for grid-structured data like images. Their core assumption is locality: the relationship between adjacent pixels in an image is far more important than that between distant pixels — a cat's ears, eyes, and whiskers are local features that, through sliding windows (convolutional kernels) extraction, are then composed layer by layer into the semantics of "cat."

Raw image → Convolutional layers (extract local features) → Pooling layers (reduce dimensions)
        → Convolutional layers → Pooling layers → … → flatten → fully connected layers → class probabilities

Two decisive advantages of CNNs over fully connected networks:

  1. Parameter sharing: one convolutional kernel slides across the entire image and reuses — for a 224×224×3 image, the first fully connected layer has tens of millions of parameters, while CNN uses only thousands;
  2. Translation invariance: whether a cat is in the top-left or bottom-right of the image, the convolutional kernel recognizes it — fully connected networks can't do this.

2. The Three Core Components of CNNs ​

1. Convolutional layers: sliding feature detectors ​

A convolutional kernel (e.g., 3×3×3) slides over the input, computing weighted sums at each position, producing a feature map. Multiple kernels = multiple feature maps (detecting different patterns: edges, colors, textures). Two hyperparameters:

  • Stride: how many pixels the kernel moves each slide; when >1, the output size shrinks;
  • Padding: zero-padding at edges to maintain output size (same padding).

Receptive field: the size of the input region that each output pixel "sees." Shallow layers have small receptive fields (looking locally), deep layers have large receptive fields (looking globally) — this is the mechanism behind "shallow layers learn edges, deep layers learn semantics."

2. Pooling layers: compression and invariance ​

Max pooling: take the maximum value in a 2×2 window — preserves the strongest response, discards exact position. Benefits: dimensionality reduction (reduces computation), introduces some translation invariance, and prevents overfitting. Modern networks (post-ResNet) tend to use "convolution with stride 2" instead of pooling, but the idea is the same.

3. Residual connections: the key to making networks deeper ​

Before ResNet in 2015, the dilemma was: as networks grew to 20+ layers, training error actually increased (degradation problem — not overfitting, but deep networks being hard to optimize). ResNet's solution is extremely simple:

output = F(x) + x     (F is a convolutional block, x is the input, i.e., the "residual")

Let the network learn the "residual" rather than the complete mapping: if a layer is already optimal, residual learning only needs to push F to 0 (output = input), sharply reducing learning difficulty. With this, ResNet reached 152 layers, ImageNet top-5 error of 3.57% — beating humans for the first time (~5.1%). Residual connections are now standard in all deep networks (they also appear in Transformers).

3. Architecture Evolution Timeline ​

YearModelKey contributionImageNet top-5 error
2012AlexNetGPU + ReLU + Dropout + data augmentation, ignited deep learning15.3% (previous year's champion: 26.2%)
2014VGGStack small kernels (3×3) deeper, clean structure7.3%
2014GoogLeNet (Inception)Multi-scale conv parallel + 1×1 conv for dimensionality reduction6.7%
2015ResNetResidual connections, 152 layers3.57% (beat humans)
2017DenseNetDense inter-layer connections, feature reuse3.46%
2019EfficientNetNeural architecture search (NAS) jointly scaling depth/width/resolutionBelow 2.9%
2020ViTTransformer eats image patches directly, start of visual foundation modelsCatching up to CNN

Trend: from "hand-designed architectures" (VGG/ResNet) to "NAS auto-search architectures" (EfficientNet), to "Transformer unifying vision and language" (ViT). Vision foundation models post-2021 (CLIP, SAM, DINO) are mostly based on ViT or CNN hybrids, but CNN's locality prior still dominates lightweight, mobile scenarios.

4. Image Task Families ​

TaskGoalRepresentative modelsMetrics
Image classificationWhole image → classResNet, EfficientNet, ViTtop-1/top-5 accuracy
Object detectionLocate + classify multiple targetsYOLO (one-stage), Faster R-CNN (two-stage)mAP
Semantic segmentationEach pixel → classU-Net, DeepLabmIoU
Instance segmentationSegment + distinguish individualsMask R-CNNmAP
Pose estimationKey point localizationHRNet, OpenPoseOKS/AP
Image retrieval / embeddingImage → vectorMetric learning, CLIPRecall@K

Detection vs. segmentation: detection gives "box + class"; segmentation gives "pixel-level mask." Business use: classification/detection for quality inspection, segmentation for medical imaging and autonomous driving, embedding for cross-modal search.

5. The Golden Recipe for Image Tasks: Transfer Learning ​

Almost all practical image projects start from pre-trained models, not training from scratch:

  1. Load pre-trained weights: ResNet/EfficientNet trained on ImageNet (free, high quality);
  2. Replace the classification head: swap the final fully connected layer to your task's output dimension;
  3. Two-stage training: first freeze the backbone and train only the new head (a few epochs), then unfreeze the backbone with a small learning rate for fine-tuning (few epochs);
  4. Data augmentation: random cropping, flipping, color jittering (feeding data with different variations each epoch, equivalent to free data expansion).
python
import torchvision.models as models
import torch.nn as nn

backbone = models.resnet18(weights=models.ResNet18_Weights.IMAGENET1K_V1)
backbone.fc = nn.Linear(512, num_classes)          # Replace classification head
# Freeze backbone (train only the classification head first)
for p in backbone.parameters():
    p.requires_grad = False
for p in backbone.fc.parameters():
    p.requires_grad = True

Why it works: low-level features (edges, textures, shapes) learned on ImageNet are universal; small dataset tasks only need to adapt high-level semantics. The freeze + fine-tune layered strategy is the core playbook for small-data image projects.

Three disciplines of transfer learning

  1. Don't overdo data augmentation: flipping/cropping destroys label semantics for "direction-sensitive" tasks (license plate recognition, medical images);
  2. Fine-tuning learning rate must be small: pre-trained weights are already good; a large learning rate ruins them in one step;
  3. Class imbalance must still be handled: pre-training solves "features," not "sample distribution."

6. Key Techniques for Training Image Models ​

  • BatchNorm: standard practice, stabilizes training, accelerates convergence, with built-in mild regularization (placed after conv, before activation);
  • Learning rate: warmup + cosine annealing, peak ~1e-3 (for fine-tuning: 1e-4~1e-5);
  • Label smoothing: soften classification head output (don't push to 100% confidence), suppresses overfitting;
  • EMA (exponential moving average): average the weights over time, use the averaged weights for inference, stable point gains;
  • Mixed precision training: FP16 forward/backward + FP32 master weights, halve VRAM, double speed (PyTorch AMP, one line to enable).

7. Visual Foundation Models and Current State (2023–2026) ​

After CNNs, vision enters the "foundation model" era:

  • CLIP (2021): jointly train image + text; the image encoder learns cross-modal semantics — "photo of a cat" and "cat" map to similar vectors, powering text-to-image (Stable Diffusion uses it for text encoding) and zero-shot classification;
  • SAM (2023): Segment Anything, prompt-driven (points/boxes/text) universal segmentation model;
  • Vision-Language Models (VLM): GPT-4V, Gemini, Qwen-VL, etc., fold vision understanding into LLMs; image QA and document understanding become mainstream;
  • Visual tokenizer trend: more systems cut images into tokens to feed to Transformers — CNN's "local prior" is being replaced by "data scale."

Current state assessment for practitioners: industry vision projects = pre-trained CNN/ViT fine-tuning + detection/segmentation frameworks (MMDetection, Detectron2) + data augmentation + distillation (large models distilled to small models for deployment). Foundation models (VLMs) handle "understanding," small models handle "high-speed inference."

8. Tradeoffs and Decision Points ​

  • CNN vs. ViT: small data, limited resources → CNN (strong prior, efficient); large data, cross-modal → ViT/visual foundation models;
  • Accuracy vs. speed: YOLO (fast) or Faster R-CNN (accurate) for detection — decide based on real-time requirements; deploy quantization + TensorRT/ONNX acceleration;
  • Fine-tuning vs. distillation: directly fine-tuning a small model is simple; distilling a large model to a small model has higher quality but heavier engineering;
  • Supervised vs. self-supervised: when labeling is expensive, use self-supervised pre-training (MAE, DINO) or CLIP zero-shot, then fine-tune with a small amount of labeled data.

Further Reading ​

References ​