Skip to content

Transformer Architecture

Quick overview The Transformer replaced recurrence with self-attention, becoming the de facto standard in the era of large models. This article breaks down the Encoder/Decoder architecture, multi-head self-attention, positional encoding, Pre-Norm/Post-Norm, provides training tricks like warmup, surveys the BERT vs GPT families and ViT/cross-modal extensions, and covers engineering optimizations like FlashAttention and KV cache.

Transformer Architecture ​

In one sentence: The Transformer is a sequence architecture built entirely on attention mechanisms, abandoning recurrence and convolution—it enables direct interaction between any two positions (path length 1) and fully parallelizable training, making it the de facto standard of the large model era (start by understanding "what attention is" in Attention Mechanisms before reading this page).

1. Why "Attention Is All You Need" ​

When Vaswani et al. introduced the Transformer in their 2017 paper Attention Is All You Need, the motivation was straightforward: the previous sequence backbone was the RNN/LSTM discussed in RNN and Sequence Modeling, which is sequential—$h_t$ depends on $h_{t-1}$, so GPU parallelism can't be utilized, and long-range dependencies are limited by gradient paths. The Transformer's response was to go "straight to the point": feed the entire sentence into self-attention at once, let every pair of positions interact directly, and have each layer's computation be independent and parallelizable. Combined with its clean, unified mathematical form, the Transformer rapidly spread from machine translation to language modeling, vision, speech, and multimodal tasks—the entire paradigm is detailed under Large Language Models (LLM).

2. Overall Architecture ​

The Transformer consists of an Encoder and Decoder, each stacked from N identical layers (the original paper used N=6), with nearly the same structure per layer:

  • Encoder layer: Multi-head self-attention → (residual + LayerNorm) → Feedforward network (FFN) → (residual + LayerNorm).
  • Decoder layer: Masked self-attention (can only attend to current and previous tokens) → (residual + LayerNorm) → Cross-attention (queries come from the decoder, keys/values come from encoder output) → (residual + LayerNorm) → FFN → (residual + LayerNorm).

Key components broken down:

  1. Self-Attention: Project inputs through three learnable matrices to obtain queries Q, keys K, and values V, then compute: $$ \text{Attention}(Q,K,V) = \text{softmax}\left(\frac{QK^\top}{\sqrt{d_k}}\right)V $$ Dividing by $\sqrt{d_k}$ prevents the variance of dot products from growing with dimension, keeping logits at reasonable magnitudes before softmax.
  2. Multi-Head: Split the d_model dimension into h heads (h=8 in the original paper), perform attention independently on each head, then concatenate—different heads learn to attend to different relationships (syntactic position, coreference, semantic similarity, etc.), like "multiple perspectives meeting in parallel."
  3. Feedforward Network (FFN): A two-layer MLP shared across positions (inner dimension expanded 4×, e.g., 512→2048→512), performing per-position nonlinear transformation. This is where most parameters accumulate.
  4. Residual + LayerNorm: Residual connections mitigate the risk of layer degradation (the same idea as ResNet in CNN and Computer Vision). LayerNorm normalizes features per token, stabilizing deep training—the difference between LayerNorm and BatchNorm is covered in Initialization and Normalization.

3. Intuition Behind Multi-Head Self-Attention ​

Why multiple heads? A single attention mechanism is "one weighted sum," carrying limited information. Multi-head allows the model to retrieve different patterns across multiple subspaces simultaneously. Intuitive analogy: at a meeting, one person looks at the big picture while sub-teams each focus on a topic (coreference, temporal order, topic), then synthesize conclusions. In practice, different heads diverge into roles like "heads that attend to neighboring words" and "heads that attend to syntactic subjects." Attention weights are often used for interpretability analysis, but be wary: they don't equal causal attribution (see Interpretability and Fairness).

4. Positional Encoding ​

Self-attention itself is permutation-invariant: "I love you" and "you love me" look identical to attention—it has no concept of "order." So positional information must be injected, through these mainstream approaches:

  • Sinusoidal positional encoding (original paper): Superimpose sine/cosine waves of different frequencies, letting the model infer relative distances from the encoding, without requiring training and allowing extrapolation to longer sequences.
  • Learnable positional embeddings: Train a table of position vectors as if positions were part of the vocabulary. BERT uses this approach.
  • RoPE (Rotary Positional Encoding, 2021): Encode positional information into the rotation matrices of Q/K, widely adopted by LLaMA, GPT-family models, and others. It has better extrapolation for relative positions.

Positional encoding is a "patch" reflecting that Transformers need structural priors—the pure structural prior is instead offloaded to the training data, which is the subtle distinction from CNNs' inductive bias (see DL Design Principles).

5. Pre-Norm vs Post-Norm ​

The combination of residual connections and LayerNorm has two arrangements, with very different training behavior:

ArrangementOrderCharacteristics
Post-NormSublayer → Residual → LayerNormOriginal paper; long gradient propagation path, unstable for deep models; requires careful initialization and warmup
Pre-NormLayerNorm → Sublayer → ResidualResidual identity path is "cleaner," more stable for deep models, faster convergence; but slightly reduced capacity, often compensated with a scale factor

Rule of thumb: small models (≤12 layers) may get slightly better accuracy with Post-Norm; deep models (GPT-family with dozens of layers) all use Pre-Norm. GPT-2, BERT, and the LLaMA series are all Pre-Norm variants. This detail reflects the engineering trade-off between "training stability vs. representational capacity."

6. Training Tricks ​

  • Learning rate warmup: Linearly ramp up the LR for the first several thousand steps, then decay at $\text{step}^{-0.5}$. Transformers are deep and attention distributions are sensitive—too high an LR at the start easily causes divergence or convergence to a bad local optimum.
  • Label smoothing: Change softmax targets from one-hot to a uniform distribution with small noise, reducing "overconfidence" and improving generalization (detailed in Overfitting and Regularization).
  • Dropout: Applied to multi-head attention outputs, FFN, and embedding layers—standard regularization for large models.
  • Initialization scaling: Initialize deeper residual branch weights with small values (e.g., 0.02)—a critical detail for making deep GPTs trainable.
  • Gradient clipping + mixed precision: Prevents explosion, saves memory, and speeds things up. The latter is covered in Training Recipes and Hyperparameter Tuning.

Engineering Tip

90% of Transformer training crashes can be traced to "attention logits exploding → softmax saturating" and "LR too high." First check the loss curve shape, then check value ranges—methods in Debugging and Diagnostics.

7. BERT (Encoder) vs GPT (Decoder) Families ​

The Transformer family splits into two camps based on "which half" is used, corresponding to two pretraining objectives:

BERT (2018, Google)GPT (2018+, OpenAI)
ArchitectureEncoder-only (bidirectional attention)Decoder-only (causal attention)
Pretraining TaskMasked Language Modeling (MLM) + Next Sentence PredictionAutoregressive "predict next token"
StrengthsUnderstanding: classification, extraction, similarityGeneration: continuation, dialogue, reasoning
DescendantsRoBERTa, DeBERTa, T5 (encoder-decoder)GPT-2/3/4, LLaMA, Qwen, DeepSeek
Current StateStill effective for understanding tasks, but overshadowed by large generative modelsThe unifying paradigm: "if it can generate, it can understand" has become the dominant belief

The logic behind the generative unification is simple: any task can be rewritten as "generate the next piece given context" (classification → "Is this review positive or negative?"). This paradigm choice profoundly shaped the subsequent Large Language Model (LLM) trajectory.

8. Extensions: ViT and Cross-Modal ​

The Transformer's generality lies in its "input-agnostic" nature—as long as inputs can be tokenized into a sequence:

  • ViT (2020): Split images into 16×16 patches, flatten them into tokens, and feed directly into the Transformer with positional encoding. Surpasses CNNs at large-scale pretraining—see CNN and Computer Vision.
  • Cross-modal Transformers: Text, images, audio, and video can all be tokenized and fed into the same attention mechanism. VLM architectures (LLaVA, Qwen-VL) that combine vision encoders with projections and LLMs, and unified multimodal models (Gemini), all stem from this—see Multimodal Models.
  • Speech: Whisper uses a Transformer for speech-to-text mapping—see Speech and Audio.

9. Complexity Optimization: FlashAttention, Sparse Attention, KV Cache ​

Transformer's O(n²) attention is a bottleneck for long sequences, and engineering has developed an entire toolkit to address it:

  • FlashAttention (2022): Tiles attention computation and performs it in SRAM, avoiding writing O(n²) intermediate matrices back to GPU memory. 2–4× speedup, linearized memory—now the standard operator for large model training and inference.
  • Sparse/Linear attention: Longformer, BigBird approximate full connectivity with local windows plus global tokens; Linformer, Performer approximate attention as low-rank or kernel forms, reducing complexity to O(n). Trade off gains against accuracy loss.
  • KV cache: During autoregressive generation, cache the K and V of historical tokens to avoid recomputing at every step—but memory grows linearly with sequence length, hence supporting optimizations like PagedAttention and speculative decoding. See the inference section in Large Language Models (LLM).

Key Insight

The evolution of Transformers is a story of "structure stabilized (2017) → scale increased (2018–2022) → efficiency revolution (2022–)." The core mathematics hasn't changed—what has changed are normalization arrangements, positional encoding, attention approximations, and engineering implementations.

10. Trade-offs ​

  • Quadratic complexity vs. linear approximation: Use sparse/linear attention for long documents and videos to save compute, but global dependencies may suffer. For general tasks, standard attention + FlashAttention remains most reliable.
  • Pre-Norm vs. Post-Norm: Stability vs. capacity. Choose Pre-Norm for deep models; for small models seeking peak accuracy, try Post-Norm.
  • Bidirectional vs. causal: Bidirectional has an information advantage for understanding tasks, but generation requires causal attention—and bidirectional BERT can't directly do generation. Use Encoder for "understanding-focused," Decoder for "generation/general-purpose."
  • Cost: Training a 13B model requires hundreds of GPUs for a month-level compute budget. Scale isn't everything—see DL Design Principles and MLOps and Model Deployment.

Further Reading ​

References ​