Theme
Transformer and NLP
Concept Definition: An Architectural Revolution that "Abandoned Recurrence"
In 2017, Vaswani et al. published "Attention Is All You Need," introducing Transformer — a sequence model based entirely on the attention mechanism that abandoned recurrence. Prior to this, the NLP king was RNN/LSTM: processing tokens one time step at a time, unable to parallelize, difficult to model long-range dependencies. Transformer allows any two tokens in a sequence to interact directly, solving both parallelization and long-range dependency problems simultaneously.
Input sequence → Embedding + position encoding → [self-attention + feed-forward network] × N → OutputWhy it matters so much: Transformer is the foundation of BERT, GPT, and all large language models (LLMs) — virtually everything "AI product" you encounter today (conversational AI, translation, search, code completion) is powered by it under the hood. It is the single most important architecture in ML over the past decade.
2. Core Mechanism: Self-Attention
Understanding attention through "looking up a dictionary"
"Attention" means: when processing a word, which words in the context should the model pay attention to? Take "it" as an example — "Xiao Ming gave the ball to Xiao Hong, and then it rolled away." "It" refers to "ball," so the model should allocate attention to "ball."
Each token generates three sets of vectors in implementation:
Q (query): "Who am I looking for?" — current token's features
K (key): "What am I?" — each token's label
V (value): "What information do I carry?" — each token's actual content
Attention weights = softmax(Q·Kᵀ / √d) — relevance between current token and each token
Output = Σ weightᵢ × Vᵢ — weighted aggregation by relevanceThe scaling factor √d prevents dot products from becoming too large, causing softmax saturation (vanishing gradients). Multi-head attention: multiple sets of Q/K/V run in parallel (e.g., 12 heads), each learning different relational subspaces (syntactic relations, coreference relations, semantic similarity), then concatenated — "looking at the sentence from multiple perspectives simultaneously."
Position encoding: making up for order
Attention itself doesn't perceive order ("cat chases dog" and "dog chases cat" have the same attention weights). Transformer injects position information using position encoding: the original paper used sine/cosine functions (fixed), while modern models use learnable positional embeddings or Rotary Positional Encoding (RoPE, the standard for LLMs).
3. Why Transformer Beat RNN
| Dimension | RNN/LSTM | Transformer |
|---|---|---|
| Parallelism | Token-by-token serial, slow training | Full sequence parallel, high GPU utilization |
| Long-range dependencies | Information decays with steps (even with LSTM) | Any two tokens reach each other in one step (O(1) path) |
| Computational complexity | O(n) steps | Self-attention O(n²) (n = sequence length) |
| Interpretability | Hidden states hard to interpret | Attention weights are observable (note: ≠ explanation, see Interpretability) |
The cost is quadratic complexity: when sequences are too long (e.g., an entire book), O(n²) explodes — this is the direction where Longformer, sparse attention, and FlashAttention have been continuously optimizing.
4. Two Paradigms: BERT vs GPT
The Transformer paper contained both an encoder (Encoder, for understanding) and a decoder (Decoder, for generation). Subsequent development diverged:
BERT route: bidirectional encoder (2018)
BERT (Bidirectional Encoder Representations from Transformers) uses bidirectional attention (each token can see both left and right context), pre-trained through two self-supervised tasks:
- Masked Language Modeling (MLM): randomly mask 15% of tokens, predict the masked ones;
- Next Sentence Prediction (NSP): determine whether two sentences are adjacent.
After pre-training, fine-tune on downstream tasks. Suited for understanding tasks: text classification, named entity recognition, question answering, semantic similarity. BERT swept 11 NLP benchmarks, establishing the "pre-training + fine-tuning" paradigm.
GPT route: autoregressive decoder (2018–)
GPT (Generative Pre-trained Transformer) uses only a decoder, unidirectional (can only see left), with the task of predicting the next token:
Input: "Today's weather is" → Predict: "really" → "good" → "."After generative pre-training, GPT can also fine-tune on understanding tasks, but GPT's core advantage is generation — it is the prototype for ChatGPT and all conversational/writing/code-generation models. The scaling law lets "predict the next word" spontaneously emerge with translation, reasoning, and programming abilities when parameters are large enough (see Large Language Models).
Comparing the two routes
| BERT (encoder) | GPT (decoder) | |
|---|---|---|
| Attention | Bidirectional | Unidirectional (causal) |
| Pre-training task | Masked prediction | Next-word prediction |
| Excels at | Understanding, classification, extraction | Generation, conversation, creation |
| Representatives | BERT, RoBERTa, DeBERTa | GPT series, Llama, Qwen |
| Scale | Hundreds of millions (0.3~400M) | Billions to trillions |
Current status: the generative route (GPT) has become the absolute mainstream due to ChatGPT's success; the BERT family remains a practical choice for low-cost understanding tasks (distilled into small models for deployment).
5. Modern Solutions for the NLP Task Family
| Task | Traditional approach | Modern approach |
|---|---|---|
| Text classification | TF-IDF + SVM/naive Bayes | Pre-trained model fine-tuning (BERT) or LLM prompting |
| Named entity recognition (NER) | CRF sequence labeling | BERT + sequence labeling head |
| Machine translation | Statistical / encoder-decoder | Large model translation (or dedicated NMT) |
| QA (extractive) | Retrieval + reading | RAG: retrieval-augmented generation |
| Text generation | Templates / LSTM | GPT-series generation + decoding strategies |
| Semantic similarity | Word embeddings + cosine | Sentence-BERT / embedding API |
| Sentiment analysis | Dictionaries / classifiers | Fine-tuning or LLM prompting |
Post-2023 paradigm shift: previously, you fine-tuned a separate BERT for each task; now the trend is to use one LLM + prompting / RAG / light fine-tuning (LoRA) to solve all tasks. Lower cost, better generalization, at the cost of more expensive inference and the need to guard against hallucinations. For choice logic, see Large Language Models (LLM) and How to Choose Frameworks and Tools.
6. Key Engineering Details of Transformer
- LayerNorm: normalizes along the feature dimension; Transformer uses LayerNorm (not BatchNorm — BatchNorm is unstable when sequence length varies);
- Residual connections: each sub-layer is
output = LayerNorm(x + Sublayer(x))— key for deep stacking; - Feed-forward network (FFN): each token passes through two fully-connected layers independently (4× hidden dim), accounting for the bulk of Transformer parameters;
- KV Cache (inference optimization): caches historical K/V during generation without re-computation, dramatically improving speed (standard for LLM inference);
- FlashAttention: memory-efficient attention implementation (IO optimization), a standard kernel for training on long sequences.
7. From NLP to Multi-modal: Attention Devours Everything
Transformer's strength lies in the fact that what the input tokens are is up to you: text is word tokens, images are cut into patches (ViT), speech is cut into frames, videos are cut into clips — the attention mechanism is modality-agnostic. This is the architectural foundation of multi-modal large models (GPT-4V, Gemini, Qwen-VL, Claude): unify images/audio/text into token streams, and one Transformer processes everything. See Large Language Models (LLM) and CNN and Computer Vision.
Three self-test questions for learning Transformer
- Can you clearly explain the roles of Q, K, V and the computation flow of attention weights?
- Can you explain why Transformer needs position encoding?
- Can you distinguish the attention differences between encoder-style (BERT) and decoder-style (GPT)? If you can answer these three smoothly, you've grasped the fundamentals of Transformer. → Classic Paper Deep Dives has a complete deep read.
8. Tradeoffs and Decision Points
- Encoder vs decoder: just need understanding/extraction → encoder (cheaper, faster); need generation/conversation → decoder (more expensive but stronger);
- Fine-tuning vs prompting: fixed task, lots of data → fine-tuning (LoRA); changing tasks, need flexibility → prompting / few-shot (saves money and time, see LLM);
- Build your own vs API: training a Transformer from scratch requires enormous compute — almost always use pre-trained models + fine-tuning, or use an API;
- Sequence length vs cost: for long text, use sparse attention / chunking / summarization; don't force it all into the window.
Further Reading
- Deep Learning Foundations — The network foundation of the attention mechanism
- Large Language Models (LLM) — The peak form of Transformer
- Generative Models — Autoregressive generation mechanisms
- CNN and Computer Vision — ViT and visual Transformer
- Classic Paper Deep Dives — Deep read of "Attention Is All You Need"
- Interpretability and Fairness — Attention ≠ explanation
- Math Primer — Mathematics of softmax and matrix multiplication
References
- Vaswani et al. Attention Is All You Need (NeurIPS 2017) — Original Transformer paper
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers (NAACL 2019)
- Radford et al. Improving Language Understanding by Generative Pre-Training (GPT-1, 2018)
- Radford et al. Language Models are Unsupervised Multitask Learners (GPT-2, 2019)
- Reimers & Gurevych. Sentence-BERT (EMNLP 2019)
- Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention (2022)
- Jurafsky & Martin. Speech and Language Processing (free online textbook) — Authoritative NLP textbook