Theme
Deep Learning Fundamentals
Concept Definition: Let Networks Learn Features on Their Own
Deep learning is a subfield of machine learning, centered on deep neural networks. Its revolution doesn't lie in "more complex networks," but in representation learning: classic ML requires hand-crafted features, while deep networks let the model automatically learn features layer by layer from raw data — shallow layers learn edges/strokes, mid-level layers learn parts/phrases, deep layers learn objects/semantics.
Classic ML: Raw data → Manual feature engineering → Shallow model → Output
Deep learning: Raw data → Deep network (auto-learns features layer by layer) → OutputA Single Neuron
The smallest unit of a neural network is the neuron (perceptron):
z = w·x + b # Weighted sum
a = σ(z) # Activation function (introduces nonlinearity)Why activation functions are necessary: without nonlinearity, a multi-layer network, no matter how deep, is equivalent to a linear model (matrix multiplication is still a linear transformation). Activation functions are the mathematical prerequisite for "deep being useful."
| Activation Function | Formula | Characteristics |
|---|---|---|
| ReLU | max(0, z) | Default choice: fast computation, mitigates vanishing gradients; downside: gradient is 0 in the negative region (dead neurons) |
| Leaky ReLU | max(0.01z, z) | Fixes the dead ReLU problem |
| GELU | Smooth ReLU variant | Standard for Transformers (used in GPT, BERT) |
| Sigmoid | 1/(1+e⁻ᶻ) | Outputs (0,1), suitable for probabilistic output; vanishing gradient in saturation zone, rarely used in hidden layers |
| Tanh | Hyperbolic tangent | Outputs (-1,1), zero-centered; usable in hidden layers, less common in deep models |
| Softmax | Multi-class normalization | Output layer only, multi-class probabilities |
The Multilayer Perceptron (MLP): Everything's Foundation
MLP = input layer + hidden layers + output layer, fully connected. Forward propagation computes the input layer by layer to the output; backpropagation uses the chain rule to propagate the loss gradient back to every layer, updating parameters with gradient descent (see Optimization and Gradient Descent).
python
import torch
import torch.nn as nn
class MLP(nn.Module):
def __init__(self, in_dim, hidden_dim, out_dim):
super().__init__()
self.net = nn.Sequential(
nn.Linear(in_dim, hidden_dim),
nn.ReLU(),
nn.Linear(hidden_dim, hidden_dim),
nn.ReLU(),
nn.Linear(hidden_dim, out_dim) # Output layer: no activation for regression, softmax for classification
)
def forward(self, x):
return self.net(x)
model = MLP(in_dim=20, hidden_dim=64, out_dim=1)The universal approximation theorem: a single-hidden-layer network with sufficient width can approximate any continuous function — but "can approximate" doesn't mean "can learn." The true value of deep networks is learning complex functions with fewer parameters and better inductive bias. That's why "deep" is more efficient than "wide."
Core Architecture Families
CNN: Images and Local Structure
Convolutional Neural Networks (CNNs) exploit the prior of image locality: adjacent pixels are highly correlated. Convolution kernels slide across the image to extract local features; parameter sharing drastically reduces the number of parameters; pooling compresses spatial dimensions.
Input image → [Conv + ReLU → Pooling] × N → Flatten → Fully connected → OutputThree key components of CNNs:
- Convolutional layers: local receptive fields + parameter sharing (one kernel reused across the whole image);
- Pooling layers: downsampling (max pooling / average pooling), dimensionality reduction + translation invariance;
- Residual connections (ResNet):
output = F(x) + x, solves the degradation problem of deep networks, the key that allows networks to reach 100+ layers.
Landmarks: LeNet (1998) → AlexNet (2012) → VGG/GoogLeNet (2014) → ResNet (2015, surpassing human) → EfficientNet (2019). See CNNs and Computer Vision.
RNN/LSTM: Sequence Modeling (the former king)
Recurrent Neural Networks (RNNs) process sequences step by step, carrying historical information in the hidden state. LSTMs use gating mechanisms (forget gate, input gate, output gate) to solve the long-term dependency and vanishing gradient problems of RNNs, dominating NLP and speech in the 2010s. Note: completely replaced by Transformers in the 2020s, but understanding the RNN "sequence state" concept remains the starting point for understanding sequence modeling (LSTMs still have a presence in some scenarios like streaming speech).
Transformer: Attention Mechanism
The 2017 paper Attention Is All You Need proposed Transformers: using self-attention to directly model dependencies between any two positions in a sequence, abandoning the recurrent structure, fully parallelizable:
Attention: Attention(Q, K, V) = softmax(QKᵀ/√d) · V- Q (query) / K (key) / V (value): each token is projected into three sets of vectors; attention weights = similarity between query and key; values are weighted by attention;
- Multi-head attention: multiple projections run in parallel, capturing different relational subspaces;
- Positional encoding: adds "order" information to parallel computation (Transformers have no inherent sense of order).
Transformer's contribution is solving both long-range dependencies and parallelism — it's the foundation of BERT/GPT and all large models; the full discussion is in Transformers and NLP.
The Complete Recipe for Training Deep Networks
Deep learning training can be summarized as a recipe table:
| Component | Default Choice | Notes |
|---|---|---|
| Data | Normalization + augmentation + shuffling | Missing normalization is the #1 cause of unstable training |
| Initialization | Xavier (sigmoid/tanh) / He (ReLU) | Wrong initialization causes gradient explosion/vanishing |
| Loss | Cross-entropy for classification / MSE or Huber for regression | Match the output distribution of the task |
| Optimizer | AdamW (Transformer) / Adam / SGD+Momentum | See Optimization |
| Learning rate | Warmup + cosine annealing, peak 1e-4~1e-3 | The most important hyperparameter |
| Regularization | Dropout + weight decay + early stopping + data augmentation | See Regularization |
| Normalization layers | BatchNorm (CNN standard) / LayerNorm (Transformer) | Critical for stable training |
The difference between normalization layers: BatchNorm normalizes along the batch dimension (depends on batch size, unstable for sequence tasks); LayerNorm normalizes along the feature dimension (independent of batch, used by Transformers). This is an easy-to-miss pitfall in engineering practice.
Deep Learning vs. Classic ML: Selection Boundaries
| Dimension | Classic ML (tree/linear) | Deep Learning |
|---|---|---|
| Data format | Primarily tabular | Images/text/speech/sequences |
| Data volume needed | Thousands to hundreds of thousands | Hundreds of thousands to billions (pre-training can mitigate) |
| Compute resources | CPU sufficient | GPU/TPU |
| Features | Manual feature engineering | Automatic representation learning |
| Interpretability | Good (trees can print rules) | Poor (black box, needs post-hoc explanation) |
| Training time | Minutes to hours | Hours to months (large models) |
Tree models often outperform deep learning on tabular data — long-standing experience from Kaggle (see Tree Models and Ensemble Learning). Deep learning isn't "better machine learning"; it's "a different approach for different problems." For the complete decision framework, see How to Choose Frameworks and Tools.
Transfer Learning and the Large Model Paradigm
One key engineering fact about deep learning: pre-trained models can be reused directly.
- Transfer learning: use ImageNet-pre-trained CNNs for feature extraction or fine-tuning, achieving good models even on small datasets — the de facto standard in computer vision;
- Pre-training + fine-tuning: BERT/GPT pre-train on massive corpora first (learning language), then fine-tune on task data — the standard paradigm for NLP;
- Prompting / in-context learning: in the large model era, "even fine-tuning isn't needed anymore" — just provide task examples (see Large Language Models).
Why transfer learning works: low-level features (edges, strokes, lexical patterns) are common across tasks; only high-level features (semantics, objects) are task-specific — so "freezing the bottom layers + fine-tuning the top layers" lets you learn good models with small data.
Tradeoffs
- Model capacity vs. data volume: when model parameters far exceed sample count, overfitting is inevitable — either add data/augmentation or reduce the model (see Regularization);
- Training cost vs. performance: the performance gain of a 10B-parameter model over the previous one may be just 1% — use hyperparameter tuning practices and efficiency techniques (mixed precision, distillation) to control costs;
- Accuracy vs. inference latency: when online requires millisecond-level response, use model compression (quantization, distillation, pruning); see MLOps;
- Reproducibility vs. exploration: deep learning experiments have high randomness (initialization, data order); fixing random seeds + experiment management is the baseline for team collaboration.
Further Reading
- Supervised Learning — neural networks are the implementers of classification/regression
- Optimization and Gradient Descent — the engine for training networks
- Overfitting and Regularization — full discussion of Dropout/early stopping
- CNNs and Computer Vision — convolution architecture in practice
- Transformers and NLP — attention mechanism in practice
- Large Language Models (LLM) — the pinnacle form of deep learning
- Math Primer for Deep Learning — linear algebra / calculus refresher
References
- Goodfellow, Bengio, Courville. Deep Learning (the "Deep Learning Bible") — the authoritative textbook on deep learning, available free online
- Nielsen. Neural Networks and Deep Learning — the clearest free book on backpropagation
- LeCun, Bengio, Hinton. Deep Learning (Nature, 2015) — a deep learning survey by three Turing Award winners
- Krizhevsky et al. ImageNet Classification with Deep CNNs (AlexNet, 2012)
- He et al. Deep Residual Learning (ResNet, 2016)
- Vaswani et al. Attention Is All You Need (2017)
- Hochreiter & Schmidhuber. Long Short-Term Memory (LSTM, 1997)