Theme
Neural Networks Fundamentals
One-line definition: A neural network is a family of "nonlinear function approximators" driven by a large number of learnable parameters — it applies "linear transformation + nonlinear activation" at each layer, and adjusts parameters so the overall function approximates the mapping we want. To answer "what is deep learning" and where it fits in the broader AI landscape, start with What is Deep Learning; this article focuses on the minimal core: "what does a network actually look like?"
I. A Single Neuron: Linear Transformation + Activation
Everything starts with the simplest artificial neuron. It does only two things:
- Linear summation: Weighted sum of each component of the input vector
x, plus a bias:
z = w·x + b = Σᵢ wᵢxᵢ + bwhere w is the weight, b is the bias, and both are learnable parameters. This step is a simple linear mapping from calculus; for linear algebra fundamentals, see Math Primer.
- Nonlinear activation: Feed
zthrough an activation functionσto produce the neuron's output:
a = σ(z)Why must step 2 exist? Because without nonlinear activation, no matter how many layers you stack, the whole network remains a linear function of the input — stacked linear transformations are equivalent to a single linear transformation, and the network's representational power collapses immediately. Nonlinearity is the prerequisite for "depth" to make sense. This is a recurring core idea throughout the Anatomy of Deep Learning Architectures.
Intuitive Analogy
A neuron is like an assembly line: Step 1 is "weighing" — assigning an importance weight to each input feature; Step 2 is a "switch" — deciding whether to "fire" the output signal based on the weighted sum. Activations like ReLU act as a switch that "lets through what exceeds the threshold, and blocks everything else."
A single neuron on its own is a linear classifier (perceptron). When Rosenblatt introduced it in 1958, it could only handle linearly separable problems. What made neural networks take off was organizing tens of thousands of neurons into layers and stacking them into depth — a scale we only needed decades later, with the computational foundation coming from backpropagation in 1986. See Backpropagation and Automatic Differentiation for details.
II. Activation Functions: Where Does Nonlinearity Come From?
The choice of activation function directly determines the network's gradient flow and representational power. The table below compares six of the most common activation functions:
| Function | Formula | Characteristics | Typical Use |
|---|---|---|---|
| ReLU | max(0, z) | Very fast to compute, gradient is constantly 1 in the positive region, mitigates vanishing gradients; but gradient is 0 in the negative region, potentially causing "dead neurons" | Default first choice, hidden layers of virtually all modern networks |
| Leaky ReLU | max(αz, z), α≈0.01 | Gives a small nonzero gradient in the negative region, reduces dead neurons | Replacement when worried about ReLU dying |
| GELU | z·Φ(z), Φ is the standard normal CDF | Smooth, probabilistic "soft gating", the standard for BERT/GPT and other Transformers | Hidden layers of Transformer-family models |
| Sigmoid | 1/(1+e⁻ᶻ) | Output in (0,1), interpretable as probability; but gradient approaches 0 in saturated regions, prone to vanishing gradients | Binary classification with probabilistic output, or gating (e.g., forget gate in LSTM) |
| Tanh | (eᶻ−e⁻ᶻ)/(eᶻ+e⁻ᶻ) | Output in (−1,1), zero-centered, symmetric; still saturates | RNN cell internals, scenarios requiring zero-centered output |
| Softmax | eᶻⁱ/Σⱼ eᶻʲ | Normalizes a set of logits into a probability distribution summing to 1 | Multi-class output layer, paired with cross-entropy |
A few rules of thumb for choosing: default ReLU in hidden layers (large models lean toward GELU), pick the output layer based on the task (Softmax for classification, typically no activation for regression), and make sure activation functions are paired with appropriate initialization and normalization — the combination of the latter two is covered in Initialization and Normalization.
Common Pitfalls
Sigmoid/Tanh are very prone to causing vanishing gradients in deep networks, because their derivatives have maximum values of only 0.25 and 1, respectively. After multiplying across many layers, gradients shrink exponentially. For deep networks, prefer ReLU-family activations; see the debugging section for troubleshooting thinking.
III. MLP: Forward Propagation and PyTorch Implementation
A Multi-Layer Perceptron (MLP) is simply neurons organized into layers: input layer → several hidden layers → output layer, with full connections within layers and unidirectional propagation between layers. The computation at layer l is:
z⁽ˡ⁾ = W⁽ˡ⁾a⁽ˡ⁻¹⁾ + b⁽ˡ⁾
a⁽ˡ⁾ = σ⁽ˡ⁾(z⁽ˡ⁾)where W⁽ˡ⁾ is the weight matrix, and a⁽⁰⁾ = x is the input. "Forward propagation" means computing this formula from layer 1 all the way through to the output layer.
A three-layer MLP in PyTorch takes just a few lines of code:
python
import torch
import torch.nn as nn
class MLP(nn.Module):
def __init__(self, in_dim, hidden_dim, out_dim):
super().__init__()
self.fc1 = nn.Linear(in_dim, hidden_dim)
self.fc2 = nn.Linear(hidden_dim, hidden_dim)
self.fc3 = nn.Linear(hidden_dim, out_dim)
self.relu = nn.ReLU()
def forward(self, x):
h = self.relu(self.fc1(x))
h = self.relu(self.fc2(h))
return self.fc3(h) # output layer typically has no activation — leave it to the loss function
model = MLP(in_dim=784, hidden_dim=256, out_dim=10)
x = torch.randn(32, 784) # one batch, 32 samples
logits = model(x) # forward propagation
assert logits.shape == (32, 10)Note the last line: the output layer returns logits (unnormalized scores), not probabilities. Passing Softmax and cross-entropy together to the loss function is numerically more stable. This is the principle of "the output layer and loss must be matched," which is exactly what Loss Functions and Output Layers covers.
Building such a network from scratch along with the training loop and watching it gradually learn is the most effective way to understand deep learning. For a hands-on path, see the "build from scratch" project in the practice section.
IV. Universal Approximation Theorem: Why "Deep Is More Efficient Than Wide"
The Universal Approximation Theorem tells us: as long as a single-hidden-layer feedforward network has enough neurons, it can approximate any continuous function to arbitrary precision.
This sounds impressive, but don't get ahead of yourself — it leads to two key implications:
- Width can approximate, but it may not be worth it. A single-hidden-layer network needs an exponential number of neurons to approximate certain functions, leading to computational explosion.
- Depth is the source of "efficiency." Splitting the same approximation capability across multiple layers, with each layer learning a "representation," allows deep networks to achieve the same accuracy with far fewer total parameters. This is why "deep is more efficient than wide": deep networks decompose complex functions into compositions of simpler, multi-level functions.
A classic intuition: simulating an XOR or sawtooth function with a wide network requires an exponential number of neurons, while a deep network can "accumulate features" layer by layer with polynomial-level parameters. More fundamentally, deep learning is not a "wider lookup table" but hierarchical representation learning — each layer abstracts the representation from the previous layer into higher-level features. This is the foundation of Representation Learning and Pretraining.
Worth Remembering
"Universal approximation" speaks to existence (it can approximate in theory), not learnability (whether gradient descent can find that function). In practice, we truly rely on depth + structural priors + large-scale data. For why structural priors matter so much, see Common Layer Types and the CNN & computer vision section.
V. Common Layer Types: An Overview
The core difference across tasks lies largely in "which layers to use." The table below lists common layer types in deep learning and their roles:
| Layer Type | Core Idea | Typical Use |
|---|---|---|
| Linear (Fully Connected) | Each output is a weighted sum of all inputs | MLPs, FFN in Transformers, classification heads |
| Conv (Convolutional) | Local receptive field + weight sharing, captures translation-invariant features | Images (see CNNs & Computer Vision) |
| Pool (Pooling) | Downsample aggregation (max/avg), reduces dimensions + increases robustness | Reducing resolution in CNNs |
| RNN/LSTM/GRU | Shared weights across time steps, modeling sequential dependencies | Sequence modeling (see the RNN & Sequence Modeling section) |
| Attention | Information aggregation weighted by relevance (QKV mechanism) | Modern sequence and image backbones (see Attention Mechanisms and Transformer Architecture) |
| Norm (Normalization) | Normalizes activations, stabilizes training | Almost every layer has one (see Initialization and Normalization) |
| Dropout | Randomly drops neurons during training, a regularization technique | Suppressing overfitting (see Overfitting and Regularization) |
An important observation: these layers are not mutually exclusive — they are combined. A typical network = several feature extraction layers (Conv/Attention) + normalization + activation + final Linear classification head. Treating "layers" as building blocks, once you understand what each block does, you can quickly decompose any modern architecture (CNN, Transformer).
VI. Intuition About Network Capacity and Depth
Capacity refers to the complexity of the function family a network can express. A naive intuition: more parameters and deeper layers mean higher capacity. But it's not always better:
- Insufficient capacity → underfitting: can't even learn on the training set, loss won't go down.
- Excessive capacity → overfitting: memorized the training set, poor generalization on test data.
- Depth vs. Width: with the same parameter count, deep networks typically learn more abstract representations and generalize better than wide networks, but they are also harder to train (gradients must propagate through more layers; see the vanishing gradient discussion in Backpropagation and Automatic Differentiation).
There is also a repeatedly validated empirical phenomenon: deep networks tend to be smoother on test sets and more robust to input perturbations than shallow networks, even with similar parameter counts. This is why ResNet (152 layers) consistently beat VGG (19 layers) — it wasn't about more parameters, but about "breaking nonlinearity across more layers," which itself produced a better solution space structure.
VII. Tradeoffs: Capacity, Data, and Structure
Tradeoffs
Capacity vs. Data: Model capacity must match data volume. A rule of thumb — when data doubles, you can increase capacity accordingly. When data is insufficient, adding parameters only accelerates overfitting. The right move is to add regularization or reduce capacity; see Overfitting and Regularization.
Width vs. Depth: Depth brings representational efficiency but increases training difficulty (vanishing gradients, optimization challenges); width brings redundancy and robustness but is less efficient. A common engineering compromise is "moderate depth + sufficient width + residual connections." Residual connections let gradients "take a shortcut" back, making them a standard in modern networks, from ResNet to Transformers without exception.
Structural Prior vs. Generality: Task-specific structures (convolution, attention) are efficient and require less data, but sacrifice generality; a fully MLP structure is the most general but hardest to train. When choosing, first ask: what kind of data am I processing, and what known invariances exist? Answers are covered in "How to Choose Frameworks and Tools" and the DL design principles section.
Choosing structure, setting capacity, and matching data -- the balance among these four pillars runs through all deep learning projects, ultimately boiled down to one sentence: start with a slightly oversized network within the bounds of your data scale, then "tame" it with regularization until it fits just right. The complete recipe is in Training Recipes and Hyperparameter Tuning.
Further Reading
- What is Deep Learning — where deep learning fits in the big picture
- Anatomy of Deep Learning Architectures — a panoramic view of the four pillars: data, models, loss, optimization
- Loss Functions and Output Layers — how output layers and losses work together
- Optimization and Gradient Descent — how parameters update once you have gradients
- Training Recipes and Hyperparameter Tuning — the engineering recipe for capacity, data, and regularization
- Attention Mechanisms — from fully connected to relevance-weighted aggregation
References
- McCulloch, Pitts. A Logical Calculus of the Ideas Immanent in Nervous Activity (1943)
- Cybenko. Approximation by Superpositions of a Sigmoidal Function (1989)
- Hornik. Approximation Capabilities of Multilayer Feedforward Networks (1991)
- Goodfellow, Bengio, Courville. Deep Learning (2016)
- PyTorch Documentation