Skip to content

Deep Learning Fundamentals

Quick overview Deep learning = deep neural networks + representation learning. This article covers everything from neurons and activation functions to forward/backward propagation, CNN/RNN/Transformer architectures, loss and training, providing a complete path to understanding deep learning from scratch, and its selection boundaries with classic ML.

Deep Learning Fundamentals ​

Concept Definition: Let Networks Learn Features on Their Own ​

Deep learning is a subfield of machine learning, centered on deep neural networks. Its revolution doesn't lie in "more complex networks," but in representation learning: classic ML requires hand-crafted features, while deep networks let the model automatically learn features layer by layer from raw data — shallow layers learn edges/strokes, mid-level layers learn parts/phrases, deep layers learn objects/semantics.

Classic ML:  Raw data → Manual feature engineering → Shallow model → Output
Deep learning: Raw data → Deep network (auto-learns features layer by layer) → Output

A Single Neuron ​

The smallest unit of a neural network is the neuron (perceptron):

z = w·x + b          # Weighted sum
a = σ(z)             # Activation function (introduces nonlinearity)

Why activation functions are necessary: without nonlinearity, a multi-layer network, no matter how deep, is equivalent to a linear model (matrix multiplication is still a linear transformation). Activation functions are the mathematical prerequisite for "deep being useful."

Activation FunctionFormulaCharacteristics
ReLUmax(0, z)Default choice: fast computation, mitigates vanishing gradients; downside: gradient is 0 in the negative region (dead neurons)
Leaky ReLUmax(0.01z, z)Fixes the dead ReLU problem
GELUSmooth ReLU variantStandard for Transformers (used in GPT, BERT)
Sigmoid1/(1+e⁻ᶻ)Outputs (0,1), suitable for probabilistic output; vanishing gradient in saturation zone, rarely used in hidden layers
TanhHyperbolic tangentOutputs (-1,1), zero-centered; usable in hidden layers, less common in deep models
SoftmaxMulti-class normalizationOutput layer only, multi-class probabilities

The Multilayer Perceptron (MLP): Everything's Foundation ​

MLP = input layer + hidden layers + output layer, fully connected. Forward propagation computes the input layer by layer to the output; backpropagation uses the chain rule to propagate the loss gradient back to every layer, updating parameters with gradient descent (see Optimization and Gradient Descent).

python
import torch
import torch.nn as nn

class MLP(nn.Module):
    def __init__(self, in_dim, hidden_dim, out_dim):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(in_dim, hidden_dim),
            nn.ReLU(),
            nn.Linear(hidden_dim, hidden_dim),
            nn.ReLU(),
            nn.Linear(hidden_dim, out_dim)   # Output layer: no activation for regression, softmax for classification
        )
    def forward(self, x):
        return self.net(x)

model = MLP(in_dim=20, hidden_dim=64, out_dim=1)

The universal approximation theorem: a single-hidden-layer network with sufficient width can approximate any continuous function — but "can approximate" doesn't mean "can learn." The true value of deep networks is learning complex functions with fewer parameters and better inductive bias. That's why "deep" is more efficient than "wide."

Core Architecture Families ​

CNN: Images and Local Structure ​

Convolutional Neural Networks (CNNs) exploit the prior of image locality: adjacent pixels are highly correlated. Convolution kernels slide across the image to extract local features; parameter sharing drastically reduces the number of parameters; pooling compresses spatial dimensions.

Input image → [Conv + ReLU → Pooling] × N → Flatten → Fully connected → Output

Three key components of CNNs:

  • Convolutional layers: local receptive fields + parameter sharing (one kernel reused across the whole image);
  • Pooling layers: downsampling (max pooling / average pooling), dimensionality reduction + translation invariance;
  • Residual connections (ResNet): output = F(x) + x, solves the degradation problem of deep networks, the key that allows networks to reach 100+ layers.

Landmarks: LeNet (1998) → AlexNet (2012) → VGG/GoogLeNet (2014) → ResNet (2015, surpassing human) → EfficientNet (2019). See CNNs and Computer Vision.

RNN/LSTM: Sequence Modeling (the former king) ​

Recurrent Neural Networks (RNNs) process sequences step by step, carrying historical information in the hidden state. LSTMs use gating mechanisms (forget gate, input gate, output gate) to solve the long-term dependency and vanishing gradient problems of RNNs, dominating NLP and speech in the 2010s. Note: completely replaced by Transformers in the 2020s, but understanding the RNN "sequence state" concept remains the starting point for understanding sequence modeling (LSTMs still have a presence in some scenarios like streaming speech).

Transformer: Attention Mechanism ​

The 2017 paper Attention Is All You Need proposed Transformers: using self-attention to directly model dependencies between any two positions in a sequence, abandoning the recurrent structure, fully parallelizable:

Attention: Attention(Q, K, V) = softmax(QKᵀ/√d) · V
  • Q (query) / K (key) / V (value): each token is projected into three sets of vectors; attention weights = similarity between query and key; values are weighted by attention;
  • Multi-head attention: multiple projections run in parallel, capturing different relational subspaces;
  • Positional encoding: adds "order" information to parallel computation (Transformers have no inherent sense of order).

Transformer's contribution is solving both long-range dependencies and parallelism — it's the foundation of BERT/GPT and all large models; the full discussion is in Transformers and NLP.

The Complete Recipe for Training Deep Networks ​

Deep learning training can be summarized as a recipe table:

ComponentDefault ChoiceNotes
DataNormalization + augmentation + shufflingMissing normalization is the #1 cause of unstable training
InitializationXavier (sigmoid/tanh) / He (ReLU)Wrong initialization causes gradient explosion/vanishing
LossCross-entropy for classification / MSE or Huber for regressionMatch the output distribution of the task
OptimizerAdamW (Transformer) / Adam / SGD+MomentumSee Optimization
Learning rateWarmup + cosine annealing, peak 1e-4~1e-3The most important hyperparameter
RegularizationDropout + weight decay + early stopping + data augmentationSee Regularization
Normalization layersBatchNorm (CNN standard) / LayerNorm (Transformer)Critical for stable training

The difference between normalization layers: BatchNorm normalizes along the batch dimension (depends on batch size, unstable for sequence tasks); LayerNorm normalizes along the feature dimension (independent of batch, used by Transformers). This is an easy-to-miss pitfall in engineering practice.

Deep Learning vs. Classic ML: Selection Boundaries ​

DimensionClassic ML (tree/linear)Deep Learning
Data formatPrimarily tabularImages/text/speech/sequences
Data volume neededThousands to hundreds of thousandsHundreds of thousands to billions (pre-training can mitigate)
Compute resourcesCPU sufficientGPU/TPU
FeaturesManual feature engineeringAutomatic representation learning
InterpretabilityGood (trees can print rules)Poor (black box, needs post-hoc explanation)
Training timeMinutes to hoursHours to months (large models)

Tree models often outperform deep learning on tabular data — long-standing experience from Kaggle (see Tree Models and Ensemble Learning). Deep learning isn't "better machine learning"; it's "a different approach for different problems." For the complete decision framework, see How to Choose Frameworks and Tools.

Transfer Learning and the Large Model Paradigm ​

One key engineering fact about deep learning: pre-trained models can be reused directly.

  • Transfer learning: use ImageNet-pre-trained CNNs for feature extraction or fine-tuning, achieving good models even on small datasets — the de facto standard in computer vision;
  • Pre-training + fine-tuning: BERT/GPT pre-train on massive corpora first (learning language), then fine-tune on task data — the standard paradigm for NLP;
  • Prompting / in-context learning: in the large model era, "even fine-tuning isn't needed anymore" — just provide task examples (see Large Language Models).

Why transfer learning works: low-level features (edges, strokes, lexical patterns) are common across tasks; only high-level features (semantics, objects) are task-specific — so "freezing the bottom layers + fine-tuning the top layers" lets you learn good models with small data.

Tradeoffs ​

  • Model capacity vs. data volume: when model parameters far exceed sample count, overfitting is inevitable — either add data/augmentation or reduce the model (see Regularization);
  • Training cost vs. performance: the performance gain of a 10B-parameter model over the previous one may be just 1% — use hyperparameter tuning practices and efficiency techniques (mixed precision, distillation) to control costs;
  • Accuracy vs. inference latency: when online requires millisecond-level response, use model compression (quantization, distillation, pruning); see MLOps;
  • Reproducibility vs. exploration: deep learning experiments have high randomness (initialization, data order); fixing random seeds + experiment management is the baseline for team collaboration.

Further Reading ​

References ​