Theme
Transformer Architecture
In one sentence: The Transformer is a sequence architecture built entirely on attention mechanisms, abandoning recurrence and convolution—it enables direct interaction between any two positions (path length 1) and fully parallelizable training, making it the de facto standard of the large model era (start by understanding "what attention is" in Attention Mechanisms before reading this page).
1. Why "Attention Is All You Need"
When Vaswani et al. introduced the Transformer in their 2017 paper Attention Is All You Need, the motivation was straightforward: the previous sequence backbone was the RNN/LSTM discussed in RNN and Sequence Modeling, which is sequential—$h_t$ depends on $h_{t-1}$, so GPU parallelism can't be utilized, and long-range dependencies are limited by gradient paths. The Transformer's response was to go "straight to the point": feed the entire sentence into self-attention at once, let every pair of positions interact directly, and have each layer's computation be independent and parallelizable. Combined with its clean, unified mathematical form, the Transformer rapidly spread from machine translation to language modeling, vision, speech, and multimodal tasks—the entire paradigm is detailed under Large Language Models (LLM).
2. Overall Architecture
The Transformer consists of an Encoder and Decoder, each stacked from N identical layers (the original paper used N=6), with nearly the same structure per layer:
- Encoder layer: Multi-head self-attention → (residual + LayerNorm) → Feedforward network (FFN) → (residual + LayerNorm).
- Decoder layer: Masked self-attention (can only attend to current and previous tokens) → (residual + LayerNorm) → Cross-attention (queries come from the decoder, keys/values come from encoder output) → (residual + LayerNorm) → FFN → (residual + LayerNorm).
Key components broken down:
- Self-Attention: Project inputs through three learnable matrices to obtain queries Q, keys K, and values V, then compute: $$ \text{Attention}(Q,K,V) = \text{softmax}\left(\frac{QK^\top}{\sqrt{d_k}}\right)V $$ Dividing by $\sqrt{d_k}$ prevents the variance of dot products from growing with dimension, keeping logits at reasonable magnitudes before softmax.
- Multi-Head: Split the d_model dimension into h heads (h=8 in the original paper), perform attention independently on each head, then concatenate—different heads learn to attend to different relationships (syntactic position, coreference, semantic similarity, etc.), like "multiple perspectives meeting in parallel."
- Feedforward Network (FFN): A two-layer MLP shared across positions (inner dimension expanded 4×, e.g., 512→2048→512), performing per-position nonlinear transformation. This is where most parameters accumulate.
- Residual + LayerNorm: Residual connections mitigate the risk of layer degradation (the same idea as ResNet in CNN and Computer Vision). LayerNorm normalizes features per token, stabilizing deep training—the difference between LayerNorm and BatchNorm is covered in Initialization and Normalization.
3. Intuition Behind Multi-Head Self-Attention
Why multiple heads? A single attention mechanism is "one weighted sum," carrying limited information. Multi-head allows the model to retrieve different patterns across multiple subspaces simultaneously. Intuitive analogy: at a meeting, one person looks at the big picture while sub-teams each focus on a topic (coreference, temporal order, topic), then synthesize conclusions. In practice, different heads diverge into roles like "heads that attend to neighboring words" and "heads that attend to syntactic subjects." Attention weights are often used for interpretability analysis, but be wary: they don't equal causal attribution (see Interpretability and Fairness).
4. Positional Encoding
Self-attention itself is permutation-invariant: "I love you" and "you love me" look identical to attention—it has no concept of "order." So positional information must be injected, through these mainstream approaches:
- Sinusoidal positional encoding (original paper): Superimpose sine/cosine waves of different frequencies, letting the model infer relative distances from the encoding, without requiring training and allowing extrapolation to longer sequences.
- Learnable positional embeddings: Train a table of position vectors as if positions were part of the vocabulary. BERT uses this approach.
- RoPE (Rotary Positional Encoding, 2021): Encode positional information into the rotation matrices of Q/K, widely adopted by LLaMA, GPT-family models, and others. It has better extrapolation for relative positions.
Positional encoding is a "patch" reflecting that Transformers need structural priors—the pure structural prior is instead offloaded to the training data, which is the subtle distinction from CNNs' inductive bias (see DL Design Principles).
5. Pre-Norm vs Post-Norm
The combination of residual connections and LayerNorm has two arrangements, with very different training behavior:
| Arrangement | Order | Characteristics |
|---|---|---|
| Post-Norm | Sublayer → Residual → LayerNorm | Original paper; long gradient propagation path, unstable for deep models; requires careful initialization and warmup |
| Pre-Norm | LayerNorm → Sublayer → Residual | Residual identity path is "cleaner," more stable for deep models, faster convergence; but slightly reduced capacity, often compensated with a scale factor |
Rule of thumb: small models (≤12 layers) may get slightly better accuracy with Post-Norm; deep models (GPT-family with dozens of layers) all use Pre-Norm. GPT-2, BERT, and the LLaMA series are all Pre-Norm variants. This detail reflects the engineering trade-off between "training stability vs. representational capacity."
6. Training Tricks
- Learning rate warmup: Linearly ramp up the LR for the first several thousand steps, then decay at $\text{step}^{-0.5}$. Transformers are deep and attention distributions are sensitive—too high an LR at the start easily causes divergence or convergence to a bad local optimum.
- Label smoothing: Change softmax targets from one-hot to a uniform distribution with small noise, reducing "overconfidence" and improving generalization (detailed in Overfitting and Regularization).
- Dropout: Applied to multi-head attention outputs, FFN, and embedding layers—standard regularization for large models.
- Initialization scaling: Initialize deeper residual branch weights with small values (e.g., 0.02)—a critical detail for making deep GPTs trainable.
- Gradient clipping + mixed precision: Prevents explosion, saves memory, and speeds things up. The latter is covered in Training Recipes and Hyperparameter Tuning.
Engineering Tip
90% of Transformer training crashes can be traced to "attention logits exploding → softmax saturating" and "LR too high." First check the loss curve shape, then check value ranges—methods in Debugging and Diagnostics.
7. BERT (Encoder) vs GPT (Decoder) Families
The Transformer family splits into two camps based on "which half" is used, corresponding to two pretraining objectives:
| BERT (2018, Google) | GPT (2018+, OpenAI) | |
|---|---|---|
| Architecture | Encoder-only (bidirectional attention) | Decoder-only (causal attention) |
| Pretraining Task | Masked Language Modeling (MLM) + Next Sentence Prediction | Autoregressive "predict next token" |
| Strengths | Understanding: classification, extraction, similarity | Generation: continuation, dialogue, reasoning |
| Descendants | RoBERTa, DeBERTa, T5 (encoder-decoder) | GPT-2/3/4, LLaMA, Qwen, DeepSeek |
| Current State | Still effective for understanding tasks, but overshadowed by large generative models | The unifying paradigm: "if it can generate, it can understand" has become the dominant belief |
The logic behind the generative unification is simple: any task can be rewritten as "generate the next piece given context" (classification → "Is this review positive or negative?"). This paradigm choice profoundly shaped the subsequent Large Language Model (LLM) trajectory.
8. Extensions: ViT and Cross-Modal
The Transformer's generality lies in its "input-agnostic" nature—as long as inputs can be tokenized into a sequence:
- ViT (2020): Split images into 16×16 patches, flatten them into tokens, and feed directly into the Transformer with positional encoding. Surpasses CNNs at large-scale pretraining—see CNN and Computer Vision.
- Cross-modal Transformers: Text, images, audio, and video can all be tokenized and fed into the same attention mechanism. VLM architectures (LLaVA, Qwen-VL) that combine vision encoders with projections and LLMs, and unified multimodal models (Gemini), all stem from this—see Multimodal Models.
- Speech: Whisper uses a Transformer for speech-to-text mapping—see Speech and Audio.
9. Complexity Optimization: FlashAttention, Sparse Attention, KV Cache
Transformer's O(n²) attention is a bottleneck for long sequences, and engineering has developed an entire toolkit to address it:
- FlashAttention (2022): Tiles attention computation and performs it in SRAM, avoiding writing O(n²) intermediate matrices back to GPU memory. 2–4× speedup, linearized memory—now the standard operator for large model training and inference.
- Sparse/Linear attention: Longformer, BigBird approximate full connectivity with local windows plus global tokens; Linformer, Performer approximate attention as low-rank or kernel forms, reducing complexity to O(n). Trade off gains against accuracy loss.
- KV cache: During autoregressive generation, cache the K and V of historical tokens to avoid recomputing at every step—but memory grows linearly with sequence length, hence supporting optimizations like PagedAttention and speculative decoding. See the inference section in Large Language Models (LLM).
Key Insight
The evolution of Transformers is a story of "structure stabilized (2017) → scale increased (2018–2022) → efficiency revolution (2022–)." The core mathematics hasn't changed—what has changed are normalization arrangements, positional encoding, attention approximations, and engineering implementations.
10. Trade-offs
- Quadratic complexity vs. linear approximation: Use sparse/linear attention for long documents and videos to save compute, but global dependencies may suffer. For general tasks, standard attention + FlashAttention remains most reliable.
- Pre-Norm vs. Post-Norm: Stability vs. capacity. Choose Pre-Norm for deep models; for small models seeking peak accuracy, try Post-Norm.
- Bidirectional vs. causal: Bidirectional has an information advantage for understanding tasks, but generation requires causal attention—and bidirectional BERT can't directly do generation. Use Encoder for "understanding-focused," Decoder for "generation/general-purpose."
- Cost: Training a 13B model requires hundreds of GPUs for a month-level compute budget. Scale isn't everything—see DL Design Principles and MLOps and Model Deployment.
Further Reading
- Large Language Models (LLMs)—All the consequences of scaling Transformers
- Attention Mechanisms—Complete derivation from definition to multi-head self-attention
- RNN and Sequence Modeling—The predecessor replaced by Transformers, and its streaming advantages
- CNN and Computer Vision—The visual battlefield of ViT
- Initialization and Normalization—LayerNorm and the engineering choice of Pre-Norm
- Training Recipes and Hyperparameter Tuning—Full panorama of recipes like warmup and label smoothing
References
- Vaswani et al. Attention Is All You Need (NeurIPS 2017)
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (NAACL 2019)
- Radford et al. Improving Language Understanding by Generative Pre-Training (2018)
- Liu et al. RoFormer: Enhanced Transformer with Rotary Position Embedding (2021)
- Dao, Fu, Ermon, Rudra, Ré. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (NeurIPS 2022)
- Xiong et al. On Layer Normalization in the Transformer Architecture (ICML 2020)
- Dosovitskiy et al. An Image is Worth 16x16 Words (ICLR 2021)
- Raffel et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) (JMLR 2020)