Appearance
Transformers and Attention
The Transformer is the sequence-model architecture proposed in the 2017 paper Attention Is All You Need, built entirely on attention, and it is the foundation of nearly every large model today.
Its core claim is compressed into the paper's title: attention is all you need. Before it, sequence modeling was ruled by recurrent neural networks (RNNs) and long short-term memory networks (LSTMs), which must process tokens one time step at a time. The Transformer threw away the recurrence and let any two positions in a sequence interact directly through attention, solving two chronic problems at once: unparallelizable training and lost long-range dependencies. Every ChatGPT, DeepSeek, Claude, and Gemini you use today runs on a Transformer underneath (almost always a Decoder-only variant); it is also the universal backbone for images (ViT), speech, multimodal models, and the U-Net inside diffusion models. To understand the Transformer is to understand the foundation of this AI era.
Where this page fits
This is a "core knowledge" page of the panoramic overview. For a quick tour of the model zoo, see Models and Leaderboards Quick Reference; to go straight to the source, head to Core Papers, Close Up.
1. Background: The RNN Bottleneck, or Why the Architecture Had to Change
To appreciate the Transformer's revolution, look at what it replaced. Until 2017, NLP was ruled by the RNN family: the model walks through tokens left to right, one time step at a time, using a hidden state to compress "everything seen so far" into a vector that gets passed to the next step.
Input: t₁ t₂ t₃ t₄
RNN: h1→ h2 → h3 → h4 # each step only sees the previous hidden state; strictly serialThis step-by-step serial design has two fatal flaws:
- The serial bottleneck: step n cannot start until step n-1 finishes, so nothing parallelizes on a GPU and training is painfully slow.
- Long-range dependencies are hard: information must survive many rounds of "compress-decompress" to travel from the start of a sentence to the end. LSTMs eased gradient vanishing with gating mechanisms (forget/input/output gates), letting information "travel further," but the problem was never cured — every step is still a bottleneck, and long-distance references and semantic links still get lost.
| Dimension | RNN | LSTM | Transformer |
|---|---|---|---|
| Parallelism | Fully serial, cannot parallelize | Fully serial, cannot parallelize | Whole sequence in parallel; high GPU utilization |
| Long-range dependencies | Poor; severe vanishing gradients | Moderate; gating helps but only partially | Any two positions connect in one step |
| Dependency path | O(n) steps | O(n) steps | O(1) steps — attention weights computed directly |
| Training cost | Slow (serial time steps) | Slow (serial time steps) | Fast (parallel) but memory-heavy (O(n²)) |
| Inference state | Relies on hidden state | Relies on hidden state | Must cache K/V (KV cache) |
| Landmark work | Elman RNN (1990) | LSTM (1997) | Transformer (2017) |
The one-line verdict
RNNs/LSTMs treat "order" as a physical process that must advance through time; the Transformer treats it as a set of relative position relations that can all be computed at once. That single leap unlocked both parallelism and long-range dependencies. A Brief History of AI traces how this technical route unfolded.
2. The Core Mechanism: Scaled Dot-Product Attention (the QKV Trio)
Understanding Attention Through a Dictionary Lookup
The intuition behind "attention": when processing a word, which words in the context should the model look at, and how much? Take this sentence:
Alice handed the ball to Bob, and then it rolled into the grass.
What does "it" refer to? Your brain instantly directs attention to "ball." That is exactly what the attention mechanism does — compute a probability distribution over all other words for the current word, then aggregate information weighted by it.
What Q, K, and V Each Mean
In implementation, every token produces three sets of vectors through three learnable projection matrices:
Q (Query) "What am I looking for?" — the "search signal" the current word sends out
K (Key) "What am I?" — the "label" attached to each word
V (Value) "What information do I carry?" — the content each word actually contributesAttention then works like this: compare Q against every K for relevance, then use those relevance scores to take a weighted sum of all the V's:
1. Relevance score = Q · Kᵀ / √d_k # dot product measures similarity; divide by √d_k to scale
2. Normalize to probabilities = softmax(scores) # weights over all positions sum to 1
3. Output = Σᵢ weightᵢ × Vᵢ # aggregate each position's information by relevanceWritten as a formula, this is Scaled Dot-Product Attention:
$$ \text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^\top}{\sqrt{d_k}}\right)V $$
The √d_k scaling matters: as the dimension d_k grows, dot-product values grow with it and push softmax into its saturated region (gradients approach 0 and learning stalls). Dividing by √d_k keeps the dot-product variance stable so that training stays stable. It is one of the most careful engineering decisions in the Transformer relative to earlier attention designs.
A Small Worked Example
Take a 2-token sequence ["cat", "mouse"] with d_k = 2. We are processing "cat," whose query vector is q = [1, 0]; the keys and values of "cat" and "mouse" are:
| token | K (key) | V (value) |
|---|---|---|
| cat | [1, 0] | [1, 1] |
| mouse | [0.5, 1] | [2, 0] |
Computing "cat"'s attention over each word:
q·k_cat = 1×1 + 0×0 = 1.0
q·k_mouse = 1×0.5 + 0×1 = 0.5
Divide by √2 → [0.707, 0.354]
softmax([0.707, 0.354]) → exp(0.707)=2.03, exp(0.354)=1.42
Weights = [2.03/(2.03+1.42), 1.42/(2.03+1.42)] = [0.59, 0.41]
Output = 0.59×[1,1] + 0.41×[2,0] = [1.41, 0.59]"Cat" attends mostly to itself (0.59) but also absorbs information from "mouse" (0.41) — semantically the two are strongly related, so even though they are not adjacent, the weight is large. This is exactly why long-range dependencies become "one step away": attention doesn't care about distance, only about content similarity.
A common misconception
A high attention weight does not mean "the model relies on this word for its explanation." The attention matrix is observable, but research has long shown that it is not the same as interpretability — remove the high-weight positions and the model usually works fine anyway. It is more accurate to think of attention as a soft alignment mechanism during training.
3. Multi-Head Attention: Looking from Several Angles at Once
A single attention head can only learn one definition of "similarity" (one set of Q/K projections). But a sentence carries several kinds of relationships at once: grammatical relations (subject-verb agreement), coreference ("it" → "ball"), semantic similarity ("cat" → "tiger"). What to do?
Multi-head attention splits attention into h "heads," each with its own independent Q/K/V projection matrices, computes h attentions in parallel, then concatenates and projects back to the original dimension:
MultiHead(Q, K, V) = Concat(head₁, head₂, …, head_h) · W_O
where headᵢ = Attention(Q·W_Qⁱ, K·W_Kⁱ, V·W_Vⁱ)Typical settings are h = 8 or 12 heads. The heads spontaneously specialize: some learn "adjacent-word dependencies," some learn "coreference," some learn "global statistics." Running many heads in parallel lets the model "view" the sequence from several relational subspaces at once, then merge the opinions. This is a core source of the Transformer's expressiveness — the paper showed that multi-head attention reliably improves results over single-head.
4. Positional Encoding: Putting "Order" Back In
Attention by itself knows nothing about order: to pure attention, "the cat chases the dog" and "the dog chases the cat" have identical word-to-word similarities. But in language, order is meaning (those two sentences mean opposite things), so position information must be injected explicitly.
The original paper used sine/cosine functions (sinusoidal positional encoding):
$$ PE_{(pos, 2i)} = \sin\left(\frac{pos}{10000^{2i/d_{model}}}\right), \quad PE_{(pos, 2i+1)} = \cos\left(\frac{pos}{10000^{2i/d_{model}}}\right) $$
where pos is the position and i is the dimension index. Sinusoidal encodings have two benefits: the mix of frequencies makes it easy for the model to learn relative positions (offsets), and there is no ceiling on sequence length.
Later mainstream options:
| Scheme | How it works | Used by |
|---|---|---|
| Sine/cosine | Generated by a fixed function; extrapolates naturally | The original Transformer paper |
| Learned positional embeddings | Position vectors learned as parameters during training | BERT, GPT-2 |
| RoPE (Rotary Position Embedding) | Encodes position into Q/K via rotation matrices; natively compatible with relative positions | Llama, Qwen, the GPT family — the modern LLM default |
| ALiBi (attention with linear biases) | No position vectors; adds distance-based biases directly to attention scores | BLOOM and others |
RoPE is now the default in nearly every mainstream large model: it folds relative-position information directly into the Q/K dot product, trains and extrapolates to longer sequences better, and pairs well with "context length extension" techniques. Why order cannot be dropped, in one line: the price of parallel computation is that order doesn't come for free — you have to encode it by hand.
5. The Standard Building Blocks: Residuals, LayerNorm, FFN
A "complete Transformer layer" is attention plus a feed-forward network, wrapped with residual connections and layer normalization (LayerNorm):
Input x
↓
x₁ = x + MultiHeadAttention(LayerNorm(x)) # attention sublayer + residual
↓
x₂ = x₁ + FFN(LayerNorm(x₁)) # feed-forward sublayer + residual
↓ Output x₂; stack N layers- Residual connections:
output = x + sublayer(x). They give gradients an "express lane" straight back to the shallow layers — the key to stacking dozens or hundreds of layers without degradation or vanishing gradients (the same idea as ResNet). - LayerNorm: normalizes along the feature dimension. Note LayerNorm, not BatchNorm — BatchNorm's statistics become unstable when sequence lengths and batch sizes keep changing, while LayerNorm does not depend on sample counts; it is the Transformer standard.
- Feed-forward network (FFN): each token independently passes through two fully connected layers (usually
4×d_modelwide with GELU activation), providing nonlinearity and memory capacity. FFN parameters account for roughly two-thirds of a model's total parameters — the large model's "knowledge warehouse."
Modern large models (Llama, for example) generally move normalization to before each sublayer (Pre-LN), which trains more stably — one of many continuing architectural refinements.
6. Three Architecture Variants: Encoder-Decoder / Encoder-only / Decoder-only
The original Attention Is All You Need architecture was an encoder-decoder: the encoder reads the full input (bidirectional attention) while the decoder generates output step by step (masked causal attention + cross-attention to the encoder's outputs). Development then split into three lines:
| Variant | Attention | Pretraining task | Good at | Representative models |
|---|---|---|---|---|
| Encoder-only | Bidirectional; the full sequence is visible | Masked language modeling (MLM) | Understanding, classification, extraction | BERT, RoBERTa |
| Decoder-only | Unidirectional (causal); only looks left | Next-token prediction | Generation, dialogue, reasoning | GPT family, Llama, Qwen, DeepSeek |
| Encoder-Decoder | Bidirectional encoder + causal decoder | Denoising / fill-in-the-blank / span corruption | Translation, summarization, transduction tasks | T5, BART, the original paper's model |
- BERT (2018) went encoder-only: mask 15% of the words at random and predict them (MLM), learn bidirectional representations, then fine-tune on downstream tasks — it swept 11 NLP benchmarks.
- GPT (2018) went decoder-only: just "predict the next token," carrying language modeling through to its logical conclusion.
- T5 (2019) went encoder-decoder: unify every task as "text-to-text," covering translation, summarization, and Q&A in one frame.
Why Decoder-only Became the Mainstream for Large Models
One observation matters most: next-token prediction is the only task that can exploit "massive unlabeled corpora + self-supervision" in full generality, and generative pretraining itself already builds understanding (the GPT paper found that while training a language model, the Transformer spontaneously develops attention "circuits" that handle Q&A, translation, and other tasks). After 2020, GPT-3 demonstrated in-context learning, showing that decoder-only models generalize better at scale and can be driven directly through prompt engineering and Retrieval-Augmented Generation (RAG).
Decoder-only is also less work engineering-wise: one model, one pretraining task, one unified generation interface, with no bidirectional encoder to maintain. Today the flagship models of OpenAI, Google, Meta, and DeepSeek are almost all decoder-only (Mistral, Llama, Qwen, and DeepSeek-R1 included) — see Large Language Models (LLMs).
The one-line verdict
Understanding-style fixed tasks on a budget → encoder models (the BERT family); dialogue / generation / general reasoning → decoder-only large models; "input-to-output" tasks like translation and summarization → encoder-decoder still holds its own.
7. Training and Inference: O(n²) Complexity and the KV Cache
Self-Attention's O(n²)
Let n be the sequence length. Computing attention means first computing QKᵀ (an n×d matrix times a d×n matrix → an n×n matrix), then multiplying by V. Every one of these matrix multiplications is O(n²), and memory usage is O(n²) as well — the Transformer's Achilles' heel:
n = 1000 → attention matrix 1000×1000 = 1 million numbers
n = 10000 → 100 million numbers (single layer, single head)
n = 100000 → 10 billion numbers (a memory disaster across layers)So "stuffing a whole book into the context" is expensive. This is exactly what inference optimization tackles (quantization, FlashAttention, KV cache quantization, and more).
The KV Cache: Why Cache K/V at Inference Time?
Large-model generation is autoregressive: emit one token at a time, append it to the context, predict the next. A naive implementation recomputes attention over the first t tokens at step t — but the K/V of earlier tokens depend only on those tokens themselves, not on the one currently being predicted. So inference engines cache every previously computed K/V in GPU memory; each step then only needs:
Existing K/V cache (the historical K and V matrices)
+ the new token's Q/K/V (only this one needs computing)
↓
Attention between Q_new and (K_cache ∪ K_new) → output the new token
↓ Append the new K/V to the cacheThis drops the per-step cost from "recompute everything, O(n²)" to "compute only the increment, O(n)" — a tens-fold throughput gain. The price is that the cache grows linearly with generation length, which is why long-form generation is memory-hungry and why techniques like KV cache quantization and GQA (grouped-query attention) exist. Practical details in Deploy and Optimize LLM Inference.
8. Influence: From NLP to Vision, Diffusion Models, and Multimodal
The Transformer's influence long ago outgrew NLP:
- Sweeping NLP: BERT (2018) and GPT (2018–) together established the two great paradigms — "pretraining + fine-tuning" and "pretraining + prompting / in-context learning" — directly leading to ChatGPT and the conversational AI era, and to reasoning models like DeepSeek-R1.
- Vision: ViT (Vision Transformer, 2020) cut images into patches and fed them to a Transformer as tokens, breaking the convolutional network's monopoly.
- Diffusion models: the U-Net backbone inside Stable Diffusion and Sora is full of attention layers — the denoising process uses attention at every resolution so image patches can "see" each other. Diffusion models borrowed this foundation from language models.
- Multimodal: cross-attention lets "the Q of text tokens" attend to "the K/V of image/audio tokens," bridging the two modalities — this is the mechanism at the heart of how GPT-4V, Gemini, and Qwen-VL "talk about pictures." See Multimodal Models.
- Recommender systems: treating a user's behavior sequence as a token stream and modeling the evolution of user interests with attention has carried into the recommender systems of the LLM era (see the recommender systems case study).
One-line summary: attention is a modality-agnostic, universal information-routing mechanism — as long as you can chop your data into a token sequence, a Transformer can process it.
9. Limits and Evolution: Life After Quadratic Complexity
The Transformer is not the end of the road; new breakthroughs keep attacking its "quadratic complexity":
| Direction | Idea | Examples |
|---|---|---|
| Sparse attention | Let only some positions interact pairwise (local windows + global tokens) | Longformer, BigBird, GPT-4's local+global pattern |
| Linear attention | Replace softmax with a factorizable kernel, turning QKᵀ into an O(n) computation | Linear Attention, Performer |
| IO-aware exact implementations | Never materialize the full n×n attention matrix; compute in tiles | FlashAttention (now the standard kernel for training and inference) |
| State-space models | Replace the KV cache with a fixed-size state; O(n) with constant inference state | Mamba, Mamba-2 |
| Hybrid architectures | Mix attention with linear/state-space layers for both expressiveness and efficiency | Jamba, Gemma 2's local attention, DeepSeek-V3's MLA |
FlashAttention is today's "unsung hero" of large-model training: through IO optimization it cuts the number of memory reads/writes during attention from O(n²) to nearly linear, sharply lowering training cost. Mamba proved that state-space models can bring sequence processing down to O(n) while holding on to quality. Even so, standard attention (paired with sparsification/hybrid strategies) remains the default choice for large models — its ability to retrieve long-range information is the most reliable. For the full frontier picture, see Frontier Progress.
A note on trade-offs
Don't assume "O(n²)" means the Transformer will be quickly replaced. In the real world, the winners are combinations of "attention first + efficiency patches (FlashAttention, GQA, MLA)"; pure linear attention still has clear weaknesses in exact long-range retrieval.
Further Reading
- Large Language Models (LLMs) — the Transformer's peak form: decoder-only at scale
- Multimodal Models — how cross-attention bridges text and images
- Diffusion Models and Generative AI — how U-Net + attention drives text-to-image
- Inference Optimization and Quantization — KV cache, GQA, and FlashAttention in practice
- Prompt Engineering — the dominant way to use a decoder-only model
- Core Papers, Close Up — a paragraph-by-paragraph reading of Attention Is All You Need
- Frontier Progress — the latest on linear attention, Mamba, and hybrid architectures
- Glossary — quick reference for attention, QKV, RoPE, and other terms
- What Are AI's Hot Concepts — the map of all core concepts
References
- Vaswani et al., Attention Is All You Need (NeurIPS 2017) — the original Transformer paper, the starting point of all discussion
- Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers (NAACL 2019) — the encoder-only line's landmark
- Radford et al., Improving Language Understanding by Generative Pre-Training (GPT-1, 2018) — the decoder-only line's origin
- Raffel et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5, 2020) — the unified encoder-decoder framework
- Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT, 2021) — the moment Transformers entered vision
- Dao et al., FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (2022) — the efficiency kernel of large-model training
- Gu & Dao, Mamba: Linear-Time Sequence Modeling with Selective State Spaces (2023) — state-space models challenge quadratic complexity
- Jay Alammar, The Illustrated Transformer — the most widely shared visual tutorial on the Transformer