Skip to content

Transformers and Attention

At a glance The 2017 Transformer — an architecture built entirely on attention and the foundation of nearly every modern large model — broken down into QKV attention, multi-head, positional encodings, the three architecture variants, O(n²) complexity, and the KV cache.

Transformers and Attention ​

The Transformer is the sequence-model architecture proposed in the 2017 paper Attention Is All You Need, built entirely on attention, and it is the foundation of nearly every large model today.

Its core claim is compressed into the paper's title: attention is all you need. Before it, sequence modeling was ruled by recurrent neural networks (RNNs) and long short-term memory networks (LSTMs), which must process tokens one time step at a time. The Transformer threw away the recurrence and let any two positions in a sequence interact directly through attention, solving two chronic problems at once: unparallelizable training and lost long-range dependencies. Every ChatGPT, DeepSeek, Claude, and Gemini you use today runs on a Transformer underneath (almost always a Decoder-only variant); it is also the universal backbone for images (ViT), speech, multimodal models, and the U-Net inside diffusion models. To understand the Transformer is to understand the foundation of this AI era.

Where this page fits

This is a "core knowledge" page of the panoramic overview. For a quick tour of the model zoo, see Models and Leaderboards Quick Reference; to go straight to the source, head to Core Papers, Close Up.

1. Background: The RNN Bottleneck, or Why the Architecture Had to Change ​

To appreciate the Transformer's revolution, look at what it replaced. Until 2017, NLP was ruled by the RNN family: the model walks through tokens left to right, one time step at a time, using a hidden state to compress "everything seen so far" into a vector that gets passed to the next step.

Input:  t₁   t₂   t₃   t₄
RNN:    h1→  h2 → h3 → h4      # each step only sees the previous hidden state; strictly serial

This step-by-step serial design has two fatal flaws:

  1. The serial bottleneck: step n cannot start until step n-1 finishes, so nothing parallelizes on a GPU and training is painfully slow.
  2. Long-range dependencies are hard: information must survive many rounds of "compress-decompress" to travel from the start of a sentence to the end. LSTMs eased gradient vanishing with gating mechanisms (forget/input/output gates), letting information "travel further," but the problem was never cured — every step is still a bottleneck, and long-distance references and semantic links still get lost.
DimensionRNNLSTMTransformer
ParallelismFully serial, cannot parallelizeFully serial, cannot parallelizeWhole sequence in parallel; high GPU utilization
Long-range dependenciesPoor; severe vanishing gradientsModerate; gating helps but only partiallyAny two positions connect in one step
Dependency pathO(n) stepsO(n) stepsO(1) steps — attention weights computed directly
Training costSlow (serial time steps)Slow (serial time steps)Fast (parallel) but memory-heavy (O(n²))
Inference stateRelies on hidden stateRelies on hidden stateMust cache K/V (KV cache)
Landmark workElman RNN (1990)LSTM (1997)Transformer (2017)

The one-line verdict

RNNs/LSTMs treat "order" as a physical process that must advance through time; the Transformer treats it as a set of relative position relations that can all be computed at once. That single leap unlocked both parallelism and long-range dependencies. A Brief History of AI traces how this technical route unfolded.

2. The Core Mechanism: Scaled Dot-Product Attention (the QKV Trio) ​

Understanding Attention Through a Dictionary Lookup ​

The intuition behind "attention": when processing a word, which words in the context should the model look at, and how much? Take this sentence:

Alice handed the ball to Bob, and then it rolled into the grass.

What does "it" refer to? Your brain instantly directs attention to "ball." That is exactly what the attention mechanism does — compute a probability distribution over all other words for the current word, then aggregate information weighted by it.

What Q, K, and V Each Mean ​

In implementation, every token produces three sets of vectors through three learnable projection matrices:

Q (Query)    "What am I looking for?"         — the "search signal" the current word sends out
K (Key)      "What am I?"                     — the "label" attached to each word
V (Value)    "What information do I carry?"   — the content each word actually contributes

Attention then works like this: compare Q against every K for relevance, then use those relevance scores to take a weighted sum of all the V's:

1. Relevance score = Q · Kᵀ / √d_k             # dot product measures similarity; divide by √d_k to scale
2. Normalize to probabilities = softmax(scores)  # weights over all positions sum to 1
3. Output = Σᵢ weightᵢ × Vᵢ                    # aggregate each position's information by relevance

Written as a formula, this is Scaled Dot-Product Attention:

$$ \text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^\top}{\sqrt{d_k}}\right)V $$

The √d_k scaling matters: as the dimension d_k grows, dot-product values grow with it and push softmax into its saturated region (gradients approach 0 and learning stalls). Dividing by √d_k keeps the dot-product variance stable so that training stays stable. It is one of the most careful engineering decisions in the Transformer relative to earlier attention designs.

A Small Worked Example ​

Take a 2-token sequence ["cat", "mouse"] with d_k = 2. We are processing "cat," whose query vector is q = [1, 0]; the keys and values of "cat" and "mouse" are:

tokenK (key)V (value)
cat[1, 0][1, 1]
mouse[0.5, 1][2, 0]

Computing "cat"'s attention over each word:

q·k_cat    = 1×1 + 0×0 = 1.0
q·k_mouse  = 1×0.5 + 0×1 = 0.5
Divide by √2 → [0.707, 0.354]
softmax([0.707, 0.354]) → exp(0.707)=2.03, exp(0.354)=1.42
Weights = [2.03/(2.03+1.42), 1.42/(2.03+1.42)] = [0.59, 0.41]
Output = 0.59×[1,1] + 0.41×[2,0] = [1.41, 0.59]

"Cat" attends mostly to itself (0.59) but also absorbs information from "mouse" (0.41) — semantically the two are strongly related, so even though they are not adjacent, the weight is large. This is exactly why long-range dependencies become "one step away": attention doesn't care about distance, only about content similarity.

A common misconception

A high attention weight does not mean "the model relies on this word for its explanation." The attention matrix is observable, but research has long shown that it is not the same as interpretability — remove the high-weight positions and the model usually works fine anyway. It is more accurate to think of attention as a soft alignment mechanism during training.

3. Multi-Head Attention: Looking from Several Angles at Once ​

A single attention head can only learn one definition of "similarity" (one set of Q/K projections). But a sentence carries several kinds of relationships at once: grammatical relations (subject-verb agreement), coreference ("it" → "ball"), semantic similarity ("cat" → "tiger"). What to do?

Multi-head attention splits attention into h "heads," each with its own independent Q/K/V projection matrices, computes h attentions in parallel, then concatenates and projects back to the original dimension:

MultiHead(Q, K, V) = Concat(head₁, head₂, …, head_h) · W_O
where headᵢ = Attention(Q·W_Qⁱ, K·W_Kⁱ, V·W_Vⁱ)

Typical settings are h = 8 or 12 heads. The heads spontaneously specialize: some learn "adjacent-word dependencies," some learn "coreference," some learn "global statistics." Running many heads in parallel lets the model "view" the sequence from several relational subspaces at once, then merge the opinions. This is a core source of the Transformer's expressiveness — the paper showed that multi-head attention reliably improves results over single-head.

4. Positional Encoding: Putting "Order" Back In ​

Attention by itself knows nothing about order: to pure attention, "the cat chases the dog" and "the dog chases the cat" have identical word-to-word similarities. But in language, order is meaning (those two sentences mean opposite things), so position information must be injected explicitly.

The original paper used sine/cosine functions (sinusoidal positional encoding):

$$ PE_{(pos, 2i)} = \sin\left(\frac{pos}{10000^{2i/d_{model}}}\right), \quad PE_{(pos, 2i+1)} = \cos\left(\frac{pos}{10000^{2i/d_{model}}}\right) $$

where pos is the position and i is the dimension index. Sinusoidal encodings have two benefits: the mix of frequencies makes it easy for the model to learn relative positions (offsets), and there is no ceiling on sequence length.

Later mainstream options:

SchemeHow it worksUsed by
Sine/cosineGenerated by a fixed function; extrapolates naturallyThe original Transformer paper
Learned positional embeddingsPosition vectors learned as parameters during trainingBERT, GPT-2
RoPE (Rotary Position Embedding)Encodes position into Q/K via rotation matrices; natively compatible with relative positionsLlama, Qwen, the GPT family — the modern LLM default
ALiBi (attention with linear biases)No position vectors; adds distance-based biases directly to attention scoresBLOOM and others

RoPE is now the default in nearly every mainstream large model: it folds relative-position information directly into the Q/K dot product, trains and extrapolates to longer sequences better, and pairs well with "context length extension" techniques. Why order cannot be dropped, in one line: the price of parallel computation is that order doesn't come for free — you have to encode it by hand.

5. The Standard Building Blocks: Residuals, LayerNorm, FFN ​

A "complete Transformer layer" is attention plus a feed-forward network, wrapped with residual connections and layer normalization (LayerNorm):

Input x
  ↓
x₁ = x + MultiHeadAttention(LayerNorm(x))     # attention sublayer + residual
  ↓
x₂ = x₁ + FFN(LayerNorm(x₁))                  # feed-forward sublayer + residual
  ↓ Output x₂; stack N layers
  • Residual connections: output = x + sublayer(x). They give gradients an "express lane" straight back to the shallow layers — the key to stacking dozens or hundreds of layers without degradation or vanishing gradients (the same idea as ResNet).
  • LayerNorm: normalizes along the feature dimension. Note LayerNorm, not BatchNorm — BatchNorm's statistics become unstable when sequence lengths and batch sizes keep changing, while LayerNorm does not depend on sample counts; it is the Transformer standard.
  • Feed-forward network (FFN): each token independently passes through two fully connected layers (usually 4×d_model wide with GELU activation), providing nonlinearity and memory capacity. FFN parameters account for roughly two-thirds of a model's total parameters — the large model's "knowledge warehouse."

Modern large models (Llama, for example) generally move normalization to before each sublayer (Pre-LN), which trains more stably — one of many continuing architectural refinements.

6. Three Architecture Variants: Encoder-Decoder / Encoder-only / Decoder-only ​

The original Attention Is All You Need architecture was an encoder-decoder: the encoder reads the full input (bidirectional attention) while the decoder generates output step by step (masked causal attention + cross-attention to the encoder's outputs). Development then split into three lines:

VariantAttentionPretraining taskGood atRepresentative models
Encoder-onlyBidirectional; the full sequence is visibleMasked language modeling (MLM)Understanding, classification, extractionBERT, RoBERTa
Decoder-onlyUnidirectional (causal); only looks leftNext-token predictionGeneration, dialogue, reasoningGPT family, Llama, Qwen, DeepSeek
Encoder-DecoderBidirectional encoder + causal decoderDenoising / fill-in-the-blank / span corruptionTranslation, summarization, transduction tasksT5, BART, the original paper's model
  • BERT (2018) went encoder-only: mask 15% of the words at random and predict them (MLM), learn bidirectional representations, then fine-tune on downstream tasks — it swept 11 NLP benchmarks.
  • GPT (2018) went decoder-only: just "predict the next token," carrying language modeling through to its logical conclusion.
  • T5 (2019) went encoder-decoder: unify every task as "text-to-text," covering translation, summarization, and Q&A in one frame.

Why Decoder-only Became the Mainstream for Large Models ​

One observation matters most: next-token prediction is the only task that can exploit "massive unlabeled corpora + self-supervision" in full generality, and generative pretraining itself already builds understanding (the GPT paper found that while training a language model, the Transformer spontaneously develops attention "circuits" that handle Q&A, translation, and other tasks). After 2020, GPT-3 demonstrated in-context learning, showing that decoder-only models generalize better at scale and can be driven directly through prompt engineering and Retrieval-Augmented Generation (RAG).

Decoder-only is also less work engineering-wise: one model, one pretraining task, one unified generation interface, with no bidirectional encoder to maintain. Today the flagship models of OpenAI, Google, Meta, and DeepSeek are almost all decoder-only (Mistral, Llama, Qwen, and DeepSeek-R1 included) — see Large Language Models (LLMs).

The one-line verdict

Understanding-style fixed tasks on a budget → encoder models (the BERT family); dialogue / generation / general reasoning → decoder-only large models; "input-to-output" tasks like translation and summarization → encoder-decoder still holds its own.

7. Training and Inference: O(n²) Complexity and the KV Cache ​

Self-Attention's O(n²) ​

Let n be the sequence length. Computing attention means first computing QKᵀ (an n×d matrix times a d×n matrix → an n×n matrix), then multiplying by V. Every one of these matrix multiplications is O(n²), and memory usage is O(n²) as well — the Transformer's Achilles' heel:

n = 1000   → attention matrix 1000×1000 = 1 million numbers
n = 10000  → 100 million numbers (single layer, single head)
n = 100000 → 10 billion numbers (a memory disaster across layers)

So "stuffing a whole book into the context" is expensive. This is exactly what inference optimization tackles (quantization, FlashAttention, KV cache quantization, and more).

The KV Cache: Why Cache K/V at Inference Time? ​

Large-model generation is autoregressive: emit one token at a time, append it to the context, predict the next. A naive implementation recomputes attention over the first t tokens at step t — but the K/V of earlier tokens depend only on those tokens themselves, not on the one currently being predicted. So inference engines cache every previously computed K/V in GPU memory; each step then only needs:

Existing K/V cache (the historical K and V matrices)
  + the new token's Q/K/V (only this one needs computing)
  ↓
Attention between Q_new and (K_cache ∪ K_new) → output the new token
↓ Append the new K/V to the cache

This drops the per-step cost from "recompute everything, O(n²)" to "compute only the increment, O(n)" — a tens-fold throughput gain. The price is that the cache grows linearly with generation length, which is why long-form generation is memory-hungry and why techniques like KV cache quantization and GQA (grouped-query attention) exist. Practical details in Deploy and Optimize LLM Inference.

8. Influence: From NLP to Vision, Diffusion Models, and Multimodal ​

The Transformer's influence long ago outgrew NLP:

  • Sweeping NLP: BERT (2018) and GPT (2018–) together established the two great paradigms — "pretraining + fine-tuning" and "pretraining + prompting / in-context learning" — directly leading to ChatGPT and the conversational AI era, and to reasoning models like DeepSeek-R1.
  • Vision: ViT (Vision Transformer, 2020) cut images into patches and fed them to a Transformer as tokens, breaking the convolutional network's monopoly.
  • Diffusion models: the U-Net backbone inside Stable Diffusion and Sora is full of attention layers — the denoising process uses attention at every resolution so image patches can "see" each other. Diffusion models borrowed this foundation from language models.
  • Multimodal: cross-attention lets "the Q of text tokens" attend to "the K/V of image/audio tokens," bridging the two modalities — this is the mechanism at the heart of how GPT-4V, Gemini, and Qwen-VL "talk about pictures." See Multimodal Models.
  • Recommender systems: treating a user's behavior sequence as a token stream and modeling the evolution of user interests with attention has carried into the recommender systems of the LLM era (see the recommender systems case study).

One-line summary: attention is a modality-agnostic, universal information-routing mechanism — as long as you can chop your data into a token sequence, a Transformer can process it.

9. Limits and Evolution: Life After Quadratic Complexity ​

The Transformer is not the end of the road; new breakthroughs keep attacking its "quadratic complexity":

DirectionIdeaExamples
Sparse attentionLet only some positions interact pairwise (local windows + global tokens)Longformer, BigBird, GPT-4's local+global pattern
Linear attentionReplace softmax with a factorizable kernel, turning QKᵀ into an O(n) computationLinear Attention, Performer
IO-aware exact implementationsNever materialize the full n×n attention matrix; compute in tilesFlashAttention (now the standard kernel for training and inference)
State-space modelsReplace the KV cache with a fixed-size state; O(n) with constant inference stateMamba, Mamba-2
Hybrid architecturesMix attention with linear/state-space layers for both expressiveness and efficiencyJamba, Gemma 2's local attention, DeepSeek-V3's MLA

FlashAttention is today's "unsung hero" of large-model training: through IO optimization it cuts the number of memory reads/writes during attention from O(n²) to nearly linear, sharply lowering training cost. Mamba proved that state-space models can bring sequence processing down to O(n) while holding on to quality. Even so, standard attention (paired with sparsification/hybrid strategies) remains the default choice for large models — its ability to retrieve long-range information is the most reliable. For the full frontier picture, see Frontier Progress.

A note on trade-offs

Don't assume "O(n²)" means the Transformer will be quickly replaced. In the real world, the winners are combinations of "attention first + efficiency patches (FlashAttention, GQA, MLA)"; pure linear attention still has clear weaknesses in exact long-range retrieval.

Further Reading ​

References ​