Theme
Core Paper Deep Dives
One-sentence summary: This page takes apart the "11 papers that changed the trajectory of LLMs" — answering five questions for each: what problem it solves (background), how it solves it (core method), why we should believe it (key experiments and numbers), where it falls short (limitations), and what to read next (extensions).
1. How to Use This Deep-Dive
- Use it as a coordinate calibrator: Locate on the paper map, choose depth from reading paths, then come back here for deep-dive.
- Read in thread order: We recommend going through 1→4→5→8→7 for the "architecture → scale → alignment" thread, then supplement with the three applications/systems papers.
- Don't stop at this page: This page is a "guided reading + key points checklist." True deep-reading means going back to the original papers. The full three-pass reading method is in Reading Discipline & FAQ.
How to use "key experiments and numbers"
Numbers go stale, but the framing of the number doesn't: GPT-2's 1.5B is parameter count, BLEU 41.8 is on WMT'14 EN-FR single model, CoT's 56.9% is GSM8K accuracy on PaLM 540B. When reading numbers, always ask "under what setup was this number achieved?"
2. Architecture Foundations (4 Papers)
1. Attention Is All You Need (2017, Vaswani et al., Google Brain)
| Dimension | Content |
|---|---|
| Background & Problem | RNN/LSTM must process sequences step by step: can't parallelize (slow training), hard to model long-range dependencies (information must "propagate" through many steps). At the time, attention was just an auxiliary plugin for RNNs. |
| Core Method | Proposed a pure-attention architecture, Transformer: ① Scaled dot-product attention (softmax(QKᵀ/√d_k)·V); ② Multi-head attention so different heads attend to different subspaces; ③ Sinusoidal/learnable positional encoding to inject order; ④ Encoder–decoder + residual connections + LayerNorm + feedforward layers. Completely eliminated recurrence and convolution. |
| Key Experiments & Numbers | Machine translation on WMT'14 EN-DE: single model BLEU 28.4, exceeding the previous best (including ensemble models) by ~2 BLEU points; EN-FR: single model BLEU 41.8, exceeding the previous best single model (41.0). Training cost: 8 P100 GPUs for 3.5 days — far less than the training time of the best systems at the time. |
| Limitations | ① Attention complexity O(n²·d) causes memory explosion for long sequences (FlashAttention three years later solved this); ② Weak positional encoding extrapolation; ③ Used encoder–decoder; GPT later proved "decoder-only" is more parameter-efficient. |
| Extensions | Transformer Architecture; Illustrated: The Illustrated Transformer; Code: The Annotated Transformer |
Deep-dive takeaways: Why divide by √d_k? Because dot products grow linearly with dimension; without scaling, softmax enters saturation and gradients vanish. √d_k keeps the variance at 1. The industry skepticism at the time was "can we really remove RNN?" — the paper answered with a 3.5-day training time and higher BLEU. Another detail: positional encoding uses sine functions, and the paper says "relative positional info can also be learned" — but later work proved that neither sine nor purely learned schemes extrapolate as well as RoPE/ALiBi, which directly leads into the entire chapter on context and long context. When reading this paper, derive the attention formula by hand (Q·Kᵀ → softmax → V) — this is the starting point for understanding everything that followed.
2. GPT-1: Improving Language Understanding by Generative Pre-Training (2018, Radford et al., OpenAI)
| Dimension | Content |
|---|---|
| Background & Problem | Before 2018, supervised learning required a large amount of labeled data; unlabeled text was everywhere but nobody used it systematically. The paper asks: can we pretrain on unsupervised text first, then fine-tune at low cost? |
| Core Method | Two-stage: ① Generative pretraining — train a Transformer decoder on unlabeled corpus using the standard language modeling objective (predict next token); ② Discriminative fine-tuning — add a linear output layer for downstream tasks. Task inputs are uniformly reformulated as token sequences (e.g., textual entailment), enabling "one architecture, many tasks." |
| Key Experiments & Numbers | 12-layer decoder, 117M parameters, pretrained on BooksCorpus; achieved SOTA at the time on 9 out of 12 NLP tasks (including natural language inference, QA, classification); auxiliary objectives designed for downstream tasks could improve an additional 1–2 points. |
| Limitations | ① Unidirectional attention: can only look left, limiting understanding tasks; ② "Pretrain + fine-tune" still required one model per task; ③ Small parameter count, no emergent abilities. |
| Extensions | Compared with BERT (read side by side) in BERT and the Encoder Family; GPT series overview in GPT Series; training data issues in Tokenization |
Deep-dive takeaways: The "generative" in the name refers to the pretraining stage using a language modeling objective, but the real legacy is the "pretrain + fine-tune" combo. The key design is unifying different tasks into token sequences (e.g., concatenating "premise + hypothesis" for entailment tasks) — the same parameters handle all tasks. This is the beginning of the "task unification" idea. Today's prompting is fundamentally "expressing tasks as text," and its bloodline traces back to this paper. Note its limitation: unidirectional context can't do true bidirectional understanding — that's BERT's breakthrough.
3. GPT-2: Language Models are Unsupervised Multitask Learners (2019, Radford et al., OpenAI)
| Dimension | Content |
|---|---|
| Background & Problem | In the fine-tuning era, every task required labeled data and a separate model. The paper pushed further: can a language model "know a bit of everything" on its own — i.e., accomplish tasks zero-shot? |
| Core Method | No task reformulation, no fine-tuning: use a bigger model + bigger data (WebText, ~40GB of text scraped from high-upvote Reddit links) for pure pretraining. Tasks are accomplished through conditioning on context via prompt. Introduced Byte-Pair Encoding (BPE) for out-of-vocabulary words. |
| Key Experiments & Numbers | 1.5B parameters, 48 layers (largest public LM at the time); zero-shot SOTA on 7 of 8 language tasks, e.g., CoQA F1 55, Summarization ROUGE 36.4; LAMBADA perplexity dropped from 99.8 to 8.6. |
| Limitations | ① Performance was clearly weaker than supervised fine-tuning (large gap remained); ② Zero-shot = "luck-based" task conditioning, unstable; ③ Generation quality good but factual quality poor (foreshadowed the later hallucination problem, see Hallucination). |
| Extensions | The few-shot route continued with GPT-3 (see Reading Paths checklist); the conditioning mechanism is expanded in Language Modeling and Inference Fundamentals |
Deep-dive takeaways: This paper was an experiment in "less is more": without any task adaptation, purely through prompt conditioning, the model naturally treats tasks as conditional probability problems. A detail worth remembering is the WebText data collection method (scraping ~40GB of text from high-upvote Reddit links) — data curation became a headline for the first time. The paper was initially withheld from open release for safety concerns, then gradually opened up — the first public discussion of "responsible release" in the LLM era. Its BPE approach for handling OOV words became the standard answer in tokenization.
4. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019, Devlin et al., Google)
| Dimension | Content |
|---|---|
| Background & Problem | GPT is unidirectional, ELMo is a shallow bidirectional concat — all "unidirectional left-looking." The paper argues: understanding tasks need true bidirectional context. |
| Core Method | Two pretraining objectives: ① Masked Language Modeling (MLM) — randomly mask 15% of tokens and have the model predict them; ② Next Sentence Prediction (NSP) — judge whether two sentences are adjacent. Use a Transformer encoder that looks both left and right simultaneously, then fine-tune for tasks. |
| Key Experiments & Numbers | BERT-Large: 340M parameters, 24 layers, 1024 hidden, 16 heads; average 80.5 on GLUE (exceeding previous SOTA by ~7 points), SQuAD v1.1 F1 93.2, SQuAD v2.0 83.1. Ablations proved: bidirectionality was the biggest gain source; NSP contributed relatively little. |
| Limitations | ① Masking only during pretraining, not at inference — train/inference mismatch; ② Encoder can't autoregressively generate, so can't do generative tasks; ③ MLM only supervises 15% of tokens per step — training-inefficient; ④ 340M parameters is small by today's standards, but "large model" in 2019. |
| Extensions | BERT family and RoBERTa/ALBERT/DistilBERT evolution in BERT and the Encoder Family; comparison with GPT route in Language Modeling |
Deep-dive takeaways: Between the two pretraining objectives, MLM is the heavy lifter; NSP was later proven nearly useless (RoBERTa without NSP performed better). The MLM masking details are worth reading closely: 15% masked tokens get 80% replaced with [MASK], 10% with random words, and 10% left unchanged — this is to mitigate the "mask at pretrain, no mask at inference" mismatch. Understanding tasks were dominated by the BERT family for nearly three years, until decoder models reclaimed SOTA on understanding tasks in 2022+ through scale and alignment. Meanwhile, today's retrieval encoders (embedding models) still heavily use bidirectional architectures — BERT's legacy didn't disappear, it just changed positions.
3. Scaling Laws (2 Papers)
5. Scaling Laws for Neural Language Models (2020, Kaplan et al., OpenAI)
| Dimension | Content |
|---|---|
| Background & Problem | Everyone intuitively believes "bigger models are better," but nobody knew: how does loss scale with parameters/data/compute as a power law? Where should money be spent — on parameters or data? |
| Core Method | Systematically scanned Transformer language models of varying scales (params ~7.7×10⁷ to 1.5×10⁹, i.e., ~77M to 1.5B), fitting three power laws: L(N) ≈ (N_c/N)^α_N, L(D) ≈ (D_c/D)^α_D, L(C) ≈ (C_c/C)^α_C. Performance is primarily controlled by the three quantities N, D, C; architecture details are secondary. Introduced a "transfer" empirical law. |
| Key Experiments & Numbers | Power law exponents α_N ≈ 0.076, α_D ≈ 0.095, α_C ≈ 0.050 (order of magnitude ~0.1); diminishing returns on params, data, and compute — but all three are "budget-transferable": when compute is multiplied by 10x, the optimal model should increase params by ~5.5x and data by 1.8x. |
| Limitations | ① Only went up to ~1B parameters; extrapolating to 175B may not hold (later proved directionally correct, but coefficients needed correction); ② Early versions underestimated the importance of data (corrected by Chinchilla); ③ Didn't cover alignment and emergent abilities — phenomena "beyond loss." |
| Extensions | Full treatment on Scaling Laws page; comparison with Chinchilla in next paper; real-world implications for training budget in Pretraining: Data and Objectives |
Deep-dive takeaways: The most counterintuitive finding was "architecture details are secondary to performance" — layer count, head count, depth-width ratio matter very little within reasonable ranges; N, D, C (the three scale quantities) are the real drivers. This gave teams a clear directive: tuning architecture is less valuable than scaling. The paper also showed "performance improves as a power law with data, but with diminishing returns," and a transfer prediction (patterns on small models can be extrapolated to thousand-times larger models). Note: the experimental window was only ~1.5B parameters — extrapolation is a "belief," not a "proof" — Chinchilla is what corrected its underestimation on the data dimension.
6. Chinchilla: Training Compute-Optimal Large Language Models (2022, Hoffmann et al., DeepMind)
| Dimension | Content |
|---|---|
| Background & Problem | The community, following Kaplan's conclusions, was blindly adding parameters (Gopher 280B, GPT-3 175B), but all were trained "data-starved." The paper challenged: under a fixed compute budget, what should the ratio of parameters to training tokens actually be? |
| Core Method | Redid the compute-optimal scan: fix total compute C, search for the optimal N and D combination. Conclusion: the optimal ratio is about 1:20 tokens per parameter (20 tokens per 1 parameter). Trained a 70B/1.4T token model (Chinchilla) per this ratio. |
| Key Experiments & Numbers | Chinchilla (70B, 1.4T tokens) outperformed Gopher (280B, 300B tokens) on average across 400+ evaluation tasks. Under the same compute, "smaller but trained longer" thoroughly beat "bigger but undertrained." Confirmed that a 70B model's loss can be comparable to a 540B model when data compensates. |
| Limitations | ① "Compute optimal" is loss-level optimal, not considering inference cost, deployment, or emergent abilities (small models aren't at a disadvantage for inference, but hard to say for capabilities); ② 1:20 is an empirical pattern — it drifts with different data quality; ③ Only answers "pretraining budget," not post-alignment training. |
| Extensions | The "add data or parameters first" practical conclusion in Scaling Laws; cost-effectiveness discussion on MoE in MoE Sparse Experts |
Deep-dive takeaways: The 1:20 derivation is worth reading carefully: fix total compute C, model loss as a function of params N and data D, F(N, D), and find the (N, D) pair that minimizes loss. The conclusion was that the community was generally "undertrained" — flagship models like GPT-3 (175B/300B tokens, ~1:1.7) and Gopher (280B/300B, ~1:1) were far from data-sufficient. Per 1:20, they should have been trained on 1.4T+ tokens. This paper's practical impact was enormous: after this, all major model training planned data according to compute-optimal ratios — "data scarcity" became a more real bottleneck than "parameter scarcity." It's also the core basis for Pretraining: Data and Objectives.
4. Alignment and Instruction (1 Paper)
7. InstructGPT: Training Language Models to Follow Instructions with Human Feedback (2022, Ouyang et al., OpenAI)
| Dimension | Content |
|---|---|
| Background & Problem | GPT-3 could generate but wasn't "obedient": didn't follow instructions, wasn't honest, could be toxic. The goal shifted to alignment — making models accountable for "helpfulness, honesty, and safety." |
| Core Method | RLHF three steps: ① SFT — supervised fine-tuning with human-written instruction-response pairs (~13K pairs); ② RM — train a reward model to score human preference comparisons (~33K preference pairs); ③ PPO — optimize policy with the RM as reward and KL penalty to constrain deviation from the initialization. |
| Key Experiments & Numbers | In human evaluation, the 1.3B InstructGPT output was preferred over the 175B GPT-3 output (win rate ~70%+); real API evaluation from user feedback was similarly better; "alignment tax" was small: average only ~0.4% drop on traditional benchmarks like SQuAD (some tasks even improved). |
| Limitations | ① Relies on human annotation — expensive, limited scale; ② Model "people-pleases" rather than "is correct" (over-refusal, sycophancy); ③ PPO training is unstable and complex to implement — this directly motivated DPO; ④ Aligns only text output, not tool use or multimodal. |
| Extensions | RLHF principles and DPO comparison in Alignment: RLHF and DPO; product story in ChatGPT and Conversational Models; alignment failure cases in Safety and Risks |
Deep-dive takeaways: The core insight is that alignment "organizes" existing capability rather than "creating" it — GPT-3 at 175B already "knew" what to do, it just didn't know how to be asked. Engineering details worth scrutinizing: SFT used only ~13K pairs (data quality matters far more than quantity); RM used ~33K human preference pairs; PPO added KL penalty to prevent the policy from deviating too far from initialization. The "alignment tax" (slight drop on traditional benchmarks) was later shown to be mitigatable through better data and multi-task training. ChatGPT is the direct product of this technology plus a conversational product — the story is in ChatGPT and Conversational Models.
5. Applications and Systems (3 Papers)
8. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020, Lewis et al., Facebook AI)
| Dimension | Content |
|---|---|
| Background & Problem | Purely parametric models "memorize" knowledge in their weights: knowledge cutoff, expensive to update, factually error-prone. The paper asks: can the model retrieve from an external corpus and make knowledge external? |
| Core Method | RAG: a retriever (DPR dense retrieval) + generator (BART-large) trained jointly. Two decoding strategies: RAG-sequence (generate full sequence per document) and RAG-token (sample tokens at each step according to document distribution). Documents are concatenated with input as conditions; end-to-end trainable. |
| Key Experiments & Numbers | Open-domain QA: RAG-sequence surpassed purely parametric T5-11B on Natural Questions; set SOTA on FEVER (fact verification) and other tasks; joint training of retriever and generator brought additional gains. |
| Limitations | ① Retrieval quality is the ceiling — if it can't retrieve, it can't generate; ② Two-stage latency is higher than pure generation; ③ Limited benefit for "open-domain chit-chat" tasks that don't need knowledge; ④ The generator can still ignore retrieval results and hallucinate — RAG isn't a panacea for hallucination (see Hallucination). |
| Extensions | Paradigm evolution (Naive/Advanced/Modular RAG) in case-studies RAG; engineering deployment in RAG in Practice; comparison with long context and fine-tuning in Context and Long Context |
Deep-dive takeaways: RAG moved knowledge storage from "inside parameters" to "outside parameters." Two design choices are worth reading carefully: the retriever uses DPR (dual-tower dense retrieval) instead of BM25 because it needs to be end-to-end differentiable; decoding uses RAG-sequence (generate a full passage per document) and RAG-token (sample at each step according to document distribution) — the former is more stable on open-domain QA, the latter is better for step-by-step citation. Its limitations foreshadowed a common challenge in all RAG systems: retrieval quality is the ceiling, and the generator may ignore retrieval results. RAG remains the primary solution for reducing hallucination (see Hallucination), though its form has evolved from naive to advanced/modular.
9. LoRA: Low-Rank Adaptation of Large Language Models (2021, Hu et al., Microsoft)
| Dimension | Content |
|---|---|
| Background & Problem | Full-parameter fine-tuning of a 175B model: each downstream task requires storing a full set of new weights (~350GB per task). GPU memory and storage are both infeasible. Can we "only learn the delta"? |
| Core Method | Freeze pretrained weights W₀, learn a low-rank delta ΔW = B·A (A, B are low-rank matrices, rank r ≪ d). Forward pass: W = W₀ + (α/r)·BA. The delta can be merged back into W₀ at inference, adding zero extra latency. |
| Key Experiments & Numbers | On GPT-3 175B, trainable parameters dropped from 175B to ~4.6M (a reduction of ~10,000×). On RoBERTa, GLUE results were comparable or better than full-parameter fine-tuning (e.g., surpassed on MRPC); rank r is insensitive within 1–64 (matrix decomposition works down to r=1). |
| Limitations | ① The low-rank assumption itself limits the representational space that can be learned (may not be enough for complex tasks); ② Hyperparameters (r, α, target_modules) require tuning; ③ The paper's scope is "adaptation," not full pretraining from scratch. |
| Extensions | Principles and hyperparams in Fine-Tuning: SFT and PEFT; end-to-end practice in Fine-Tuning Practice: Full LoRA Workflow; QLoRA quantized version in Frontier Trends |
Deep-dive takeaways: The low-rank assumption has experimental backing: performing SVD on updates to pretrained weights reveals that the dominant singular values are few — the delta does live in a low-rank subspace. A practical conclusion worth remembering: rank r between 1 and 64 makes almost no performance difference (experiments on GPT-3 in the paper), so "higher r is more stable" is a common misconception; α controls scaling, and in practice α/r is often kept constant. LoRA's significance isn't just saving GPU memory — it made "one base + multiple adapters" feasible, directly spawning the entire parameter-efficient fine-tuning family (QLoRA, DoRA, etc.) and forming the theoretical foundation for Fine-Tuning: SFT and PEFT.
10. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (2022, Dao et al., Stanford)
| Dimension | Content |
|---|---|
| Background & Problem | Attention O(n²) memory: at n=64k, a single fp32 attention matrix (64k × 64k) requires ~16GB. The bottleneck isn't just compute — it's memory bandwidth (HBM) — standard implementations write S and P back to HBM and read them back. |
| Core Method | I/O-aware exact attention: ① Tiling — partition Q, K, V into blocks and compute on SRAM; ② Recomputation — recompute the attention matrix during backpropagation, avoiding O(n²) intermediate storage; ③ Use an online update trick for exact softmax to maintain bit-accurate consistency with standard attention. |
| Key Experiments & Numbers | Memory reduced from O(n²) to O(n). Compared to standard PyTorch attention, end-to-end training speedup: BERT-large ~15%, GPT-2 ~3×, Long-Range Arena ~2.4× (attention operator itself up to ~7.6×, increasing with sequence length). |
| Limitations | ① Only optimizes "standard attention" form; not directly applicable to sparse/linear attention; ② Engineering implementation depends on GPU architecture details; ③ Doesn't reduce attention compute complexity (still O(n²) compute). |
| Extensions | Attention complexity and long-context engineering challenges in Context and Long Context; KV Cache optimization at inference time in Inference Fundamentals; FlashAttention-2 and PagedAttention in Frontier Trends |
Deep-dive takeaways: The disruptive insight is "the bottleneck is I/O, not FLOPs": standard implementations write intermediate matrices S, P back to HBM and read them back, but HBM bandwidth is far slower than on-chip SRAM. Two technical pillars: tiling (block so SRAM can hold it) and recompute (recompute attention during backprop to save O(n²) intermediate storage). It maintains "exact" attention (numerically identical to standard implementation) — not approximate — so it can be drop-in-replaced into any attention. FA-2 further optimizes parallelization and tiling for another ~2× speedup. Reading this alongside understanding KV Cache connects the dots to Inference Fundamentals.
11. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022, Wei et al., Google)
| Dimension | Content |
|---|---|
| Background & Problem | Large models perform poorly on math and multi-step reasoning: asking "what's the answer?" directly yields intuitive mistakes. The paper found: letting the model "think before answering" significantly elicits reasoning ability. |
| Core Method | Chain-of-thought prompting: provide "intermediate reasoning steps" in few-shot examples (few-shot CoT); or simply append "Let's think step by step" (the prototype for zero-shot CoT). Reasoning steps aren't trained — they're elicited by prompting. |
| Key Experiments & Numbers | On PaLM 540B, GSM8K (elementary math) accuracy went from ~17.9% to ~56.9%; dramatic improvement on multi-subject exam benchmarks too. Scale threshold: models below ~100B get little to no CoT benefit — sometimes negative — classic evidence of "emergent abilities." |
| Limitations | ① Sensitive to model scale — ineffective for small models; ② CoT can be "confidently wrong" — reasoning process looks reasonable but conclusion is wrong; ③ Gains are most pronounced for tasks requiring rigor (e.g., math), limited for commonsense reasoning. |
| Extensions | CoT and self-consistency engineering in Prompting and Prompting Practice; emergent ability debates in Scaling Laws |
Deep-dive takeaways: "Emergence" is the most controversial and important observation in this paper: CoT shows no benefit (or even harm) below ~100B parameters, then suddenly improves significantly above it. This suggests reasoning ability isn't directly taught by the training objective — it's "revealed" once scale is sufficient. The way few-shot CoT examples are written matters — they should present a complete "step-by-step thinking" process. Later, zero-shot CoT ("Let's think step by step") was proven even simpler. Its successors: self-consistency uses multiple sampling + majority vote for further improvement; Tree of Thoughts (ToT) expands the search space. Note the boundary: CoT improves the "process" not "conclusion correctness" — cases where reasoning seems reasonable but the conclusion is wrong are common.
6. Cross-Paper Narrative: One Thread
These 11 papers together tell the complete LLM story — "decoder + scale + alignment":
text
Attention (2017) gave the "parallelizable, long-range" skeleton
→ GPT-1/2 (2018–2019) established the "decoder + pretraining" route
→ Scaling Laws (2020) gave the theoretical permission for "adding parameters"
→ GPT-3 empirical proof (2020) spread the scale belief (deep-dive in [Reading Paths](/papers/paths))
→ InstructGPT (2022) turned "can generate" into "can obey"
→ Chinchilla (2022) corrected the ratio: data and params are equally important
→ RAG/LoRA/CoT/FlashAttention (2020–2022) solved "affordable, usable, fast"Three comparative relationships worth remembering:
| Comparison | Takeaway |
|---|---|
| GPT (unidirectional decoder) vs BERT (bidirectional encoder) | Route split between generation and understanding tasks; decoder ultimately won |
| Scaling Laws (add parameters) vs Chinchilla (add data) | Under the same compute, params to tokens at ~1:20 is optimal |
| RLHF (InstructGPT, train reward model) vs Prompting (CoT, elicit reasoning) | One changes weights, one changes context — they're stackable |
7. Recommended Reading Order for the 11 Papers
The 11 papers don't need to be read sequentially by number. This grouping is more efficient (numbers correspond to sections on this page):
| Stage | Recommended Group | Reason |
|---|---|---|
| Step 1 | 1 (Attention) + 2 (GPT-1) | Build the architecture skeleton first |
| Step 2 | 4 (BERT) + 3 (GPT-2) | Two routes face each other like mirrors |
| Step 3 | 5 (Scaling) + 6 (Chinchilla) | Scale thread: belief first, then correction |
| Step 4 | 7 (InstructGPT) | The turning point from scale to alignment |
| Step 5 | 11 (CoT) + 8 (RAG) | Applications: prompting and retrieval |
| Step 6 | 9 (LoRA) + 10 (FlashAttention) | Systems: fine-tuning and acceleration |
Within a group, read comparatively (e.g., 1 and 10: "the same attention, from architecture to system"). Between groups, order is not interchangeable: scale without architecture is a house of cards, alignment without scale has nothing to align.
8. Citation Relationships Between Papers
The value of deep-reading is seeing "who stands on whose shoulders":
| Paper | Stands on the shoulders of | Inspired |
|---|---|---|
| Attention | Bahdanau attention, Seq2Seq | Almost all subsequent papers |
| GPT-1 | Attention (removed encoder) | GPT-2/3, the entire decoder family |
| BERT | Attention + ELMo | RoBERTa, T5, the encoder family |
| Scaling Laws | Scale observations from GPT-2/3 | Chinchilla, all training budget planning |
| Chinchilla | Scaling Laws | Modern training ratio standards |
| InstructGPT | RLHF (2017) + GPT-3 | ChatGPT, DPO, Constitutional AI |
| LoRA | Adapter (2019) | QLoRA, PEFT family |
| FlashAttention | Algorithm + systems crossover | FA-2, PagedAttention |
| CoT | GPT-3 few-shot | Self-consistency, ToT, o1 |
| RAG | DPR, BART | Advanced RAG, multimodal retrieval |
A starting point for following the vine
After completing this table, you have 10 "citation chains." From any paper, click into its citations, and you enter the expansion mode of the paper map — the map is alive.
9. Common Follow-Up Questions in Interviews
These 11 papers nearly cover all the high-frequency topics for ML interviews (full question bank in Interview Questions):
| Paper | Possible Interview Follow-Up | Key Answer Points |
|---|---|---|
| Attention | Why divide by √d_k? Why do multiple heads work? | Prevent softmax saturation; heads attend to different subspaces |
| GPT-2 | What's the difference between zero-shot and few-shot? | Prompt is the task description; few-shot adds examples |
| BERT | Is NSP necessary? | RoBERTa proved it can be removed; MLM is the real driver |
| Scaling Laws | Add data or add parameters? | Check your current N:D ratio; reference 1:20 |
| Chinchilla | What is compute-optimal? | Find optimal (N, D) pair under fixed compute |
| InstructGPT | What does each of the RLHF three steps do? | SFT → RM → PPO, KL penalty prevents drift |
| LoRA | How to set r and α? | r between 1–64 makes almost no difference, keep α/r constant |
| FlashAttention | Why does it save memory? | I/O-aware: tiling + recomputation |
| CoT | Why doesn't it work for small models? | Capability emergence has a scale threshold |
| RAG | What if retrieval fails? | Hybrid retrieval, query rewriting, reranking |
10. What to Do After Deep-Reading
- Read the originals: This page is a navigation map — cross-reference each section with the original (links in the references below).
- Expand along citation chains: Continue from the "extensions" column on this page + each paper's Related Work.
- Take notes: Use the card note method to build a five-section card per paper: "problem / method / numbers / limitations / my scenario."
- Connect to the frontier: These are pre-2022 foundations. For 2023–2025 developments, see Frontier Trends.
Further Reading
- Reading Paths — The complete checklist and reading methods corresponding to these deep-dives
- Paper Map — Place the 11 papers back on the timeline coordinate system
- Transformer Architecture — Companion concept page for deep-dive paper 1
- Scaling Laws — Companion concept pages for deep-dives 5 and 6
- Alignment: RLHF and DPO — Companion concept page for deep-dive 7
- Frontier Trends — Where to go after finishing the classics
References
Official originals for the 11 deep-dive papers (real arXiv links):
- Attention Is All You Need (2017)
- Radford et al. Improving Language Understanding by Generative Pre-Training (GPT-1, 2018) — GPT-1 original paper (OpenAI official PDF; not formally published on arXiv)
- Language Models are Unsupervised Multitask Learners / GPT-2 (2019)
- BERT: Pre-training of Deep Bidirectional Transformers (2019)
- Scaling Laws for Neural Language Models (2020)
- Training Compute-Optimal Large Language Models / Chinchilla (2022)
- Training Language Models to Follow Instructions with Human Feedback / InstructGPT (2022)
- LoRA: Low-Rank Adaptation of Large Language Models (2021)
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (2022)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020)
- Deep Reinforcement Learning from Human Preferences (2017) — RLHF origin, prerequisite reading for deep-dive 7