Skip to content

Transformer and NLP

Quick overview The 2017 paper "Attention Is All You Need" dominated natural language processing with an attention mechanism. This article dissects the Transformer architecture (self-attention/multi-head/position encoding), the encoder-decoder paradigm, the divergence between BERT and GPT, and modern solutions for the NLP task family.

Transformer and NLP ​

Concept Definition: An Architectural Revolution that "Abandoned Recurrence" ​

In 2017, Vaswani et al. published "Attention Is All You Need," introducing Transformer — a sequence model based entirely on the attention mechanism that abandoned recurrence. Prior to this, the NLP king was RNN/LSTM: processing tokens one time step at a time, unable to parallelize, difficult to model long-range dependencies. Transformer allows any two tokens in a sequence to interact directly, solving both parallelization and long-range dependency problems simultaneously.

Input sequence → Embedding + position encoding → [self-attention + feed-forward network] × N → Output

Why it matters so much: Transformer is the foundation of BERT, GPT, and all large language models (LLMs) — virtually everything "AI product" you encounter today (conversational AI, translation, search, code completion) is powered by it under the hood. It is the single most important architecture in ML over the past decade.

2. Core Mechanism: Self-Attention ​

Understanding attention through "looking up a dictionary" ​

"Attention" means: when processing a word, which words in the context should the model pay attention to? Take "it" as an example — "Xiao Ming gave the ball to Xiao Hong, and then it rolled away." "It" refers to "ball," so the model should allocate attention to "ball."

Each token generates three sets of vectors in implementation:

Q (query):        "Who am I looking for?"    — current token's features
K (key):          "What am I?"               — each token's label
V (value):        "What information do I carry?" — each token's actual content

Attention weights = softmax(Q·Kᵀ / √d)   — relevance between current token and each token
Output = Σ weightᵢ × Vᵢ               — weighted aggregation by relevance

The scaling factor √d prevents dot products from becoming too large, causing softmax saturation (vanishing gradients). Multi-head attention: multiple sets of Q/K/V run in parallel (e.g., 12 heads), each learning different relational subspaces (syntactic relations, coreference relations, semantic similarity), then concatenated — "looking at the sentence from multiple perspectives simultaneously."

Position encoding: making up for order ​

Attention itself doesn't perceive order ("cat chases dog" and "dog chases cat" have the same attention weights). Transformer injects position information using position encoding: the original paper used sine/cosine functions (fixed), while modern models use learnable positional embeddings or Rotary Positional Encoding (RoPE, the standard for LLMs).

3. Why Transformer Beat RNN ​

DimensionRNN/LSTMTransformer
ParallelismToken-by-token serial, slow trainingFull sequence parallel, high GPU utilization
Long-range dependenciesInformation decays with steps (even with LSTM)Any two tokens reach each other in one step (O(1) path)
Computational complexityO(n) stepsSelf-attention O(n²) (n = sequence length)
InterpretabilityHidden states hard to interpretAttention weights are observable (note: ≠ explanation, see Interpretability)

The cost is quadratic complexity: when sequences are too long (e.g., an entire book), O(n²) explodes — this is the direction where Longformer, sparse attention, and FlashAttention have been continuously optimizing.

4. Two Paradigms: BERT vs GPT ​

The Transformer paper contained both an encoder (Encoder, for understanding) and a decoder (Decoder, for generation). Subsequent development diverged:

BERT route: bidirectional encoder (2018) ​

BERT (Bidirectional Encoder Representations from Transformers) uses bidirectional attention (each token can see both left and right context), pre-trained through two self-supervised tasks:

  • Masked Language Modeling (MLM): randomly mask 15% of tokens, predict the masked ones;
  • Next Sentence Prediction (NSP): determine whether two sentences are adjacent.

After pre-training, fine-tune on downstream tasks. Suited for understanding tasks: text classification, named entity recognition, question answering, semantic similarity. BERT swept 11 NLP benchmarks, establishing the "pre-training + fine-tuning" paradigm.

GPT route: autoregressive decoder (2018–) ​

GPT (Generative Pre-trained Transformer) uses only a decoder, unidirectional (can only see left), with the task of predicting the next token:

Input: "Today's weather is"  →  Predict: "really" → "good" → "."

After generative pre-training, GPT can also fine-tune on understanding tasks, but GPT's core advantage is generation — it is the prototype for ChatGPT and all conversational/writing/code-generation models. The scaling law lets "predict the next word" spontaneously emerge with translation, reasoning, and programming abilities when parameters are large enough (see Large Language Models).

Comparing the two routes ​

BERT (encoder)GPT (decoder)
AttentionBidirectionalUnidirectional (causal)
Pre-training taskMasked predictionNext-word prediction
Excels atUnderstanding, classification, extractionGeneration, conversation, creation
RepresentativesBERT, RoBERTa, DeBERTaGPT series, Llama, Qwen
ScaleHundreds of millions (0.3~400M)Billions to trillions

Current status: the generative route (GPT) has become the absolute mainstream due to ChatGPT's success; the BERT family remains a practical choice for low-cost understanding tasks (distilled into small models for deployment).

5. Modern Solutions for the NLP Task Family ​

TaskTraditional approachModern approach
Text classificationTF-IDF + SVM/naive BayesPre-trained model fine-tuning (BERT) or LLM prompting
Named entity recognition (NER)CRF sequence labelingBERT + sequence labeling head
Machine translationStatistical / encoder-decoderLarge model translation (or dedicated NMT)
QA (extractive)Retrieval + readingRAG: retrieval-augmented generation
Text generationTemplates / LSTMGPT-series generation + decoding strategies
Semantic similarityWord embeddings + cosineSentence-BERT / embedding API
Sentiment analysisDictionaries / classifiersFine-tuning or LLM prompting

Post-2023 paradigm shift: previously, you fine-tuned a separate BERT for each task; now the trend is to use one LLM + prompting / RAG / light fine-tuning (LoRA) to solve all tasks. Lower cost, better generalization, at the cost of more expensive inference and the need to guard against hallucinations. For choice logic, see Large Language Models (LLM) and How to Choose Frameworks and Tools.

6. Key Engineering Details of Transformer ​

  • LayerNorm: normalizes along the feature dimension; Transformer uses LayerNorm (not BatchNorm — BatchNorm is unstable when sequence length varies);
  • Residual connections: each sub-layer is output = LayerNorm(x + Sublayer(x)) — key for deep stacking;
  • Feed-forward network (FFN): each token passes through two fully-connected layers independently (4× hidden dim), accounting for the bulk of Transformer parameters;
  • KV Cache (inference optimization): caches historical K/V during generation without re-computation, dramatically improving speed (standard for LLM inference);
  • FlashAttention: memory-efficient attention implementation (IO optimization), a standard kernel for training on long sequences.

7. From NLP to Multi-modal: Attention Devours Everything ​

Transformer's strength lies in the fact that what the input tokens are is up to you: text is word tokens, images are cut into patches (ViT), speech is cut into frames, videos are cut into clips — the attention mechanism is modality-agnostic. This is the architectural foundation of multi-modal large models (GPT-4V, Gemini, Qwen-VL, Claude): unify images/audio/text into token streams, and one Transformer processes everything. See Large Language Models (LLM) and CNN and Computer Vision.

Three self-test questions for learning Transformer

  1. Can you clearly explain the roles of Q, K, V and the computation flow of attention weights?
  2. Can you explain why Transformer needs position encoding?
  3. Can you distinguish the attention differences between encoder-style (BERT) and decoder-style (GPT)? If you can answer these three smoothly, you've grasped the fundamentals of Transformer. → Classic Paper Deep Dives has a complete deep read.

8. Tradeoffs and Decision Points ​

  • Encoder vs decoder: just need understanding/extraction → encoder (cheaper, faster); need generation/conversation → decoder (more expensive but stronger);
  • Fine-tuning vs prompting: fixed task, lots of data → fine-tuning (LoRA); changing tasks, need flexibility → prompting / few-shot (saves money and time, see LLM);
  • Build your own vs API: training a Transformer from scratch requires enormous compute — almost always use pre-trained models + fine-tuning, or use an API;
  • Sequence length vs cost: for long text, use sparse attention / chunking / summarization; don't force it all into the window.

Further Reading ​

References ​