Skip to content

CNN and Computer Vision

Quick overview Convolutional Neural Networks were the first breakthrough that ignited deep learning. This article breaks down the three core building blocks—convolution, pooling, and residual connections—along with the architectural evolution from LeNet to ViT, covers the family of vision tasks, the transfer learning recipe (with PyTorch code) and training tricks, and surveys the current state of large vision models like CLIP, SAM, and VLMs.

CNN and Computer Vision ​

In one sentence: Convolutional Neural Networks (CNNs) are deep models that bake in image priors—locality, weight sharing, and translation equivariance—through three core building blocks: convolution, pooling, and residual connections. They were the first major architecture to ignite deep learning in 2012, and remain the underlying engine for virtually all vision systems today (for the broader timeline, start with A Brief History of Deep Learning's Evolution).

1. Why Images Need CNNs ​

Consider a counterexample: take a 224×224×3 color image, flatten it into ~150K numbers, and feed it into a fully connected layer, as described in our Neural Networks Primer. If the first layer has 1024 neurons, that's 150K × 1024 ≈ 150M parameters—the weight matrix itself is larger than the input image. This reveals three problems with applying fully connected layers to images:

  1. Explosion of parameters: With many pixels and deep layers, parameter counts become astronomical. Training demands both data and compute that are impractical.
  2. Loss of structure: Flattening destroys the 2D adjacency of pixels—the spatial structure of "cat on the left, background on the right" disappears.
  3. No translation invariance: Shift the same cat by 10 pixels, and a fully connected network sees a completely different input, forcing it to learn this from data alone.

CNNs solve this by injecting "image priors" into the architecture: local connections (each neuron sees only a small neighborhood), weight sharing (the same filter slides across the entire image), and hierarchical features (shallow layers learn edges and textures, deeper layers learn parts and semantics). This inductive bias dramatically shrinks the hypothesis space, enabling the model to learn generalizable visual representations with far fewer parameters.

2. The Three Building Blocks ​

Convolutional Layer ​

A convolution kernel (filter) is a small matrix, e.g., 3×3, that slides over the input feature map performing element-wise multiply-and-add to produce an output feature map. The output dimensions are determined by three hyperparameters: stride (how many pixels the kernel moves each step), padding (zero-padding at the borders), and number of kernels (determines the output channels). For example, with a 224×224×3 input, 64 kernels of size 3×3, stride 1, and padding 1, the output is 224×224×64—same spatial resolution, just different channels. The output at position (i, j) is:

$$ y_{i,j,k} = b_k + \sum_{c=0}^{C-1}\sum_{p=0}^{K-1}\sum_{q=0}^{K-1} w_{p,q,k,c}\cdot x_{i+p,,j+q,,c} $$

The parameter count for one convolution = K×K×input channels×output channels (plus biases). For example, 3×3×256×256 ≈ 590K—versus 256²×256² for a fully connected layer with the same 256-channel input and output. 1×1 convolutions don't aggregate spatial information; they only mix channels, commonly used for dimensionality reduction or as "learnable channel-wise weighting." You'll see this idea in many subsequent architectures, including the MLP blocks in Transformers.

Pooling Layer ​

Pooling downsamples a neighborhood: max pooling takes the local maximum (preserving the strongest response and providing some translation invariance), average pooling takes the mean (smoother). A 2×2 pool halves the height and width of the feature map, has zero parameters, and serves three purposes: reducing computation, enlarging the receptive field of subsequent convolutions, and providing mild invariance. Pooling's role has diminished in modern architectures (ResNet replaces it with strided convolutions), but understanding it is still the starting point for grasping the CNN feature pyramid.

Residual Connections ​

Residual connections are what make CNNs able to go "deep." Intuitively: if several layers can't learn a useful transformation, the safest choice is to "do nothing" (identity mapping). By packaging layers as residual blocks, the learning target becomes the residual F(x) = H(x) − x, with output H(x) = F(x) + x. This addition leaves a "highway" for gradients: during backpropagation, gradients can flow through the skip connections with minimal loss (see Backpropagation and Automatic Differentiation for chain rule details), mitigating the vanishing gradient problem in deep networks. ResNet used this structure to build a 152-layer network and simultaneously solved the degradation problem (the phenomenon where deeper networks perform worse).

Intuitive Analogy

Convolution "looks locally," pooling "looks more broadly," and residuals "go deeper safely"—together they form a coarse-grained simulation of the visual cortex that "understands better, sees more, and stays stable" as it processes.

3. Architectural Evolution ​

The table below summarizes the key milestones of the CNN lineage. ImageNet Top-5 error rates are the official published results from the ILSVRC for single-model or ensemble models (Top-1 classification is the representative metric):

ModelYearKey ContributionImageNet Top-5 Error
LeNet-51998Classic combo of convolution + pooling + FC, handwritten digit recognition— (MNIST, not ILSVRC)
AlexNet2012ReLU, Dropout, data augmentation, dual GPU; the deep learning inflection point15.3% (winner, far ahead of 2nd place at 26.2%)
VGG2014Stacking small 3×3 kernels instead of large ones, same receptive field, fewer parameters7.3%
GoogLeNet (Inception)2014Inception multi-branch parallelism + 1×1 convolutions for dimensionality reduction6.67% (winner)
ResNet2015Residual connections, 152-layer training feasible3.57% (winner, surpassing human-level 5.1% for the first time)
DenseNet2017Dense inter-layer connections, feature reuse, shorter gradient paths5.19% (competitive that year)
EfficientNet2019Neural architecture search + compound scaling (depth/width/resolution co-optimized)2.27% (SOTA at the time)
ViT2020Pure Transformer structure for images, outperforms CNNs with large-scale pretraining88.55% Top-1 accuracy (pretrained on ImageNet-21k)

Three patterns emerge: architectures are getting deeper and wider (LeNet 5 layers → ResNet 152 layers), inductive biases are gradually giving way to data and compute (ViT has no convolution prior, making up for it with massive pretraining data), and automation participates in design (EfficientNet uses NAS to search for structure). For the deeper trade-off between "structural priors vs. data scale," see DL Design Principles.

4. The Family of Vision Tasks ​

Image tasks go far beyond classification, but classification is often the foundation (the feature extractor of a classification network is reused by virtually all downstream tasks):

TaskGoalRepresentative Methods
Image ClassificationOne class per imageAlexNet, ResNet, EfficientNet, ViT
Object DetectionLocalize + classify multiple objectsR-CNN family (two-stage), YOLO/SSD (one-stage), DETR (Transformer)
Semantic SegmentationOne class per pixelFCN, U-Net, DeepLab
Instance SegmentationDistinguish each independent object instanceMask R-CNN
Pose EstimationKey point localizationOpenPose, HRNet
Image RetrievalFind similar images by contentMetric learning + vector retrieval (feature index)

Detection and segmentation both build on the pipeline of "extract features, then localize/classify"; pose estimation is essentially dense keypoint regression. Public datasets like COCO and PASCAL VOC are commonly used for these tasks; see Datasets and Tools Reference for details. An engineering reality: these tasks are rarely trained from scratch—almost all start from ImageNet-pretrained backbone networks, which leads us to the next practical section.

5. The Transfer Learning Recipe ​

One of the most practical recipes on the entire site. The core idea comes from Representation Learning and Pretraining: first learn "generic visual features" on massive general-purpose data, then adapt with a small amount of task-specific data. The classic four-step flow:

  1. Pretrain: Train the backbone on large corpora like ImageNet (14 million images), learning hierarchical features from edges → textures → parts → semantics;
  2. Replace the head: Swap the final fully connected classification layer of the pretrained network with a new head (change the number of classes to your target);
  3. Freeze and fine-tune with layer-wise learning rates: First freeze the backbone (set learning_rate to 0 or use a learning rate one to two orders of magnitude lower), train only the new head; when you have enough data, gradually unfreeze higher layers;
  4. Data augmentation: Random cropping, flipping, color jittering (ColorJitter), CutMix, etc.—effectively free data expansion.

PyTorch reference implementation (ResNet50 based on torchvision):

python
import torch
import torchvision.models as models

model = models.resnet50(weights=models.ResNet50_Weights.IMAGENET1K_V2)
# Replace the head: change from 1000 classes to 10 target classes
model.fc = torch.nn.Linear(model.fc.in_features, 10)

# Freeze the backbone (freeze top layers as needed)
for name, param in model.named_parameters():
    if "fc" not in name:
        param.requires_grad = False

# Layer-wise fine-tuning: low LR for backbone, normal LR for the new head
optimizer = torch.optim.AdamW([
    {"params": model.fc.parameters(), "lr": 1e-3},
    {"params": [p for n, p in model.named_parameters()
                if "fc" not in n and p.requires_grad], "lr": 1e-4},
])

This recipe works across classification, detection, segmentation, and retrieval. Details on fine-tuning hyperparameters (when to unfreeze, what learning rates to use) are in Training Recipes and Hyperparameter Tuning.

6. Training Tricks ​

Six high-frequency tips for training CNNs from scratch or fine-tuning:

  • BatchNorm: Normalizes each channel per batch, allowing higher learning rates, accelerating convergence, and providing mild regularization. Training and inference behave differently (inference uses running statistics); details in Initialization and Normalization.
  • Learning rate scheduling: Linear warmup followed by cosine annealing is the current default; within an epoch, EMA (exponential moving average of weights) often gives a free 0.3–0.5% accuracy boost.
  • Label smoothing: Replaces the 1 in one-hot with 0.9 / 0.1×uniform distribution, reducing overconfidence and improving generalization (Szegedy et al., 2016)—it's a form of "output regularization" under Overfitting and Regularization.
  • Mixed precision (AMP): FP16/FP32 mixed computation with loss scaling halves memory usage and doubles speed—now standard practice.
  • Data augmentation: Base augmentations plus strong ones like CutMix/MixUp significantly improve robustness.
  • Stochastic depth / pruning: Randomly drop residual blocks during training for more stable deep network training.

Common Pitfalls

ResNet's default hyperparameters on ImageNet (batch size 256, initial lr 0.1, 90 epochs) don't transfer directly to small datasets or different tasks. Before copying someone else's configuration, think through the data scale—otherwise overfitting or underfitting is likely. For debugging strategies, see Debugging and Diagnostics.

7. The State of Large Vision Models ​

The next step for CNNs is "from specialist to generalist." After 2021, vision models rapidly moved toward multimodal and foundation model paradigms:

  • CLIP (OpenAI, 2021): Contrastive learning on 400M image-text pairs, aligning image and text features into a shared space, enabling zero-shot image classification and image-text retrieval.
  • SAM (Segment Anything, Meta, 2023): A prompting segmentation foundation model trained on 11M images + 1B masks, capable of segmenting any object.
  • VLMs (Vision-Language Models): GPT-4V, LLaVA, Qwen-VL, etc. connect vision encoders to large language models, enabling the model to "describe images," answer visual questions, and perform reasoning—see Multimodal Models for details on these architectures.

These models share a common paradigm: pretrain on large-scale weakly supervised contrastive or generative data, unifying vision, text (and even audio) into a single representation space. Traditional CNNs often degrade to a "vision encoder" component in these systems, with the backbone shifting to Transformer architecture.

8. Trade-offs ​

  • Inductive bias vs. data efficiency: CNNs win on small data thanks to priors; ViT wins at massive scale thanks to data and compute. If your data is under a million images, CNNs or hybrid architectures are often more stable; at billion-scale corpora, pure attention structures have a higher ceiling.
  • Depth vs. interpretability: Deeper is more accurate, but features become harder to interpret. "Why is this picture classified as a cat?" is often a black box. See Interpretability and Fairness for related discussions.
  • Accuracy vs. speed: EfficientNet's NAS finds the "accuracy/compute Pareto frontier." Actual deployment still requires engineering optimizations like pruning, quantization, and distillation—see MLOps and Model Deployment.
  • Convergence of vision and language paradigms: Self-attention has penetrated vision; convolutions have been introduced into Transformers (e.g., ConvNeXt's "modernized ResNet"). The boundary between the two is blurring—no need to pick a side.

Further Reading ​

References ​