Skip to content

Core Paper Deep Dives

At a glance In-depth readings of 11 classic papers that shaped the LLM landscape: Attention Is All You Need, GPT-1/2, BERT, Scaling Laws, Chinchilla, InstructGPT, LoRA, FlashAttention, CoT, RAG — each with background problem, core method, key experimental numbers, limitations, and extensions.

Core Paper Deep Dives ​

One-sentence summary: This page takes apart the "11 papers that changed the trajectory of LLMs" — answering five questions for each: what problem it solves (background), how it solves it (core method), why we should believe it (key experiments and numbers), where it falls short (limitations), and what to read next (extensions).

1. How to Use This Deep-Dive ​

  1. Use it as a coordinate calibrator: Locate on the paper map, choose depth from reading paths, then come back here for deep-dive.
  2. Read in thread order: We recommend going through 1→4→5→8→7 for the "architecture → scale → alignment" thread, then supplement with the three applications/systems papers.
  3. Don't stop at this page: This page is a "guided reading + key points checklist." True deep-reading means going back to the original papers. The full three-pass reading method is in Reading Discipline & FAQ.

How to use "key experiments and numbers"

Numbers go stale, but the framing of the number doesn't: GPT-2's 1.5B is parameter count, BLEU 41.8 is on WMT'14 EN-FR single model, CoT's 56.9% is GSM8K accuracy on PaLM 540B. When reading numbers, always ask "under what setup was this number achieved?"

2. Architecture Foundations (4 Papers) ​

1. Attention Is All You Need (2017, Vaswani et al., Google Brain) ​

DimensionContent
Background & ProblemRNN/LSTM must process sequences step by step: can't parallelize (slow training), hard to model long-range dependencies (information must "propagate" through many steps). At the time, attention was just an auxiliary plugin for RNNs.
Core MethodProposed a pure-attention architecture, Transformer: ① Scaled dot-product attention (softmax(QKᵀ/√d_k)·V); ② Multi-head attention so different heads attend to different subspaces; ③ Sinusoidal/learnable positional encoding to inject order; ④ Encoder–decoder + residual connections + LayerNorm + feedforward layers. Completely eliminated recurrence and convolution.
Key Experiments & NumbersMachine translation on WMT'14 EN-DE: single model BLEU 28.4, exceeding the previous best (including ensemble models) by ~2 BLEU points; EN-FR: single model BLEU 41.8, exceeding the previous best single model (41.0). Training cost: 8 P100 GPUs for 3.5 days — far less than the training time of the best systems at the time.
Limitations① Attention complexity O(n²·d) causes memory explosion for long sequences (FlashAttention three years later solved this); ② Weak positional encoding extrapolation; ③ Used encoder–decoder; GPT later proved "decoder-only" is more parameter-efficient.
ExtensionsTransformer Architecture; Illustrated: The Illustrated Transformer; Code: The Annotated Transformer

Deep-dive takeaways: Why divide by √d_k? Because dot products grow linearly with dimension; without scaling, softmax enters saturation and gradients vanish. √d_k keeps the variance at 1. The industry skepticism at the time was "can we really remove RNN?" — the paper answered with a 3.5-day training time and higher BLEU. Another detail: positional encoding uses sine functions, and the paper says "relative positional info can also be learned" — but later work proved that neither sine nor purely learned schemes extrapolate as well as RoPE/ALiBi, which directly leads into the entire chapter on context and long context. When reading this paper, derive the attention formula by hand (Q·Kᵀ → softmax → V) — this is the starting point for understanding everything that followed.

2. GPT-1: Improving Language Understanding by Generative Pre-Training (2018, Radford et al., OpenAI) ​

DimensionContent
Background & ProblemBefore 2018, supervised learning required a large amount of labeled data; unlabeled text was everywhere but nobody used it systematically. The paper asks: can we pretrain on unsupervised text first, then fine-tune at low cost?
Core MethodTwo-stage: ① Generative pretraining — train a Transformer decoder on unlabeled corpus using the standard language modeling objective (predict next token); ② Discriminative fine-tuning — add a linear output layer for downstream tasks. Task inputs are uniformly reformulated as token sequences (e.g., textual entailment), enabling "one architecture, many tasks."
Key Experiments & Numbers12-layer decoder, 117M parameters, pretrained on BooksCorpus; achieved SOTA at the time on 9 out of 12 NLP tasks (including natural language inference, QA, classification); auxiliary objectives designed for downstream tasks could improve an additional 1–2 points.
Limitations① Unidirectional attention: can only look left, limiting understanding tasks; ② "Pretrain + fine-tune" still required one model per task; ③ Small parameter count, no emergent abilities.
ExtensionsCompared with BERT (read side by side) in BERT and the Encoder Family; GPT series overview in GPT Series; training data issues in Tokenization

Deep-dive takeaways: The "generative" in the name refers to the pretraining stage using a language modeling objective, but the real legacy is the "pretrain + fine-tune" combo. The key design is unifying different tasks into token sequences (e.g., concatenating "premise + hypothesis" for entailment tasks) — the same parameters handle all tasks. This is the beginning of the "task unification" idea. Today's prompting is fundamentally "expressing tasks as text," and its bloodline traces back to this paper. Note its limitation: unidirectional context can't do true bidirectional understanding — that's BERT's breakthrough.

3. GPT-2: Language Models are Unsupervised Multitask Learners (2019, Radford et al., OpenAI) ​

DimensionContent
Background & ProblemIn the fine-tuning era, every task required labeled data and a separate model. The paper pushed further: can a language model "know a bit of everything" on its own — i.e., accomplish tasks zero-shot?
Core MethodNo task reformulation, no fine-tuning: use a bigger model + bigger data (WebText, ~40GB of text scraped from high-upvote Reddit links) for pure pretraining. Tasks are accomplished through conditioning on context via prompt. Introduced Byte-Pair Encoding (BPE) for out-of-vocabulary words.
Key Experiments & Numbers1.5B parameters, 48 layers (largest public LM at the time); zero-shot SOTA on 7 of 8 language tasks, e.g., CoQA F1 55, Summarization ROUGE 36.4; LAMBADA perplexity dropped from 99.8 to 8.6.
Limitations① Performance was clearly weaker than supervised fine-tuning (large gap remained); ② Zero-shot = "luck-based" task conditioning, unstable; ③ Generation quality good but factual quality poor (foreshadowed the later hallucination problem, see Hallucination).
ExtensionsThe few-shot route continued with GPT-3 (see Reading Paths checklist); the conditioning mechanism is expanded in Language Modeling and Inference Fundamentals

Deep-dive takeaways: This paper was an experiment in "less is more": without any task adaptation, purely through prompt conditioning, the model naturally treats tasks as conditional probability problems. A detail worth remembering is the WebText data collection method (scraping ~40GB of text from high-upvote Reddit links) — data curation became a headline for the first time. The paper was initially withheld from open release for safety concerns, then gradually opened up — the first public discussion of "responsible release" in the LLM era. Its BPE approach for handling OOV words became the standard answer in tokenization.

4. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019, Devlin et al., Google) ​

DimensionContent
Background & ProblemGPT is unidirectional, ELMo is a shallow bidirectional concat — all "unidirectional left-looking." The paper argues: understanding tasks need true bidirectional context.
Core MethodTwo pretraining objectives: ① Masked Language Modeling (MLM) — randomly mask 15% of tokens and have the model predict them; ② Next Sentence Prediction (NSP) — judge whether two sentences are adjacent. Use a Transformer encoder that looks both left and right simultaneously, then fine-tune for tasks.
Key Experiments & NumbersBERT-Large: 340M parameters, 24 layers, 1024 hidden, 16 heads; average 80.5 on GLUE (exceeding previous SOTA by ~7 points), SQuAD v1.1 F1 93.2, SQuAD v2.0 83.1. Ablations proved: bidirectionality was the biggest gain source; NSP contributed relatively little.
Limitations① Masking only during pretraining, not at inference — train/inference mismatch; ② Encoder can't autoregressively generate, so can't do generative tasks; ③ MLM only supervises 15% of tokens per step — training-inefficient; ④ 340M parameters is small by today's standards, but "large model" in 2019.
ExtensionsBERT family and RoBERTa/ALBERT/DistilBERT evolution in BERT and the Encoder Family; comparison with GPT route in Language Modeling

Deep-dive takeaways: Between the two pretraining objectives, MLM is the heavy lifter; NSP was later proven nearly useless (RoBERTa without NSP performed better). The MLM masking details are worth reading closely: 15% masked tokens get 80% replaced with [MASK], 10% with random words, and 10% left unchanged — this is to mitigate the "mask at pretrain, no mask at inference" mismatch. Understanding tasks were dominated by the BERT family for nearly three years, until decoder models reclaimed SOTA on understanding tasks in 2022+ through scale and alignment. Meanwhile, today's retrieval encoders (embedding models) still heavily use bidirectional architectures — BERT's legacy didn't disappear, it just changed positions.

3. Scaling Laws (2 Papers) ​

5. Scaling Laws for Neural Language Models (2020, Kaplan et al., OpenAI) ​

DimensionContent
Background & ProblemEveryone intuitively believes "bigger models are better," but nobody knew: how does loss scale with parameters/data/compute as a power law? Where should money be spent — on parameters or data?
Core MethodSystematically scanned Transformer language models of varying scales (params ~7.7×10⁷ to 1.5×10⁹, i.e., ~77M to 1.5B), fitting three power laws: L(N) ≈ (N_c/N)^α_N, L(D) ≈ (D_c/D)^α_D, L(C) ≈ (C_c/C)^α_C. Performance is primarily controlled by the three quantities N, D, C; architecture details are secondary. Introduced a "transfer" empirical law.
Key Experiments & NumbersPower law exponents α_N ≈ 0.076, α_D ≈ 0.095, α_C ≈ 0.050 (order of magnitude ~0.1); diminishing returns on params, data, and compute — but all three are "budget-transferable": when compute is multiplied by 10x, the optimal model should increase params by ~5.5x and data by 1.8x.
Limitations① Only went up to ~1B parameters; extrapolating to 175B may not hold (later proved directionally correct, but coefficients needed correction); ② Early versions underestimated the importance of data (corrected by Chinchilla); ③ Didn't cover alignment and emergent abilities — phenomena "beyond loss."
ExtensionsFull treatment on Scaling Laws page; comparison with Chinchilla in next paper; real-world implications for training budget in Pretraining: Data and Objectives

Deep-dive takeaways: The most counterintuitive finding was "architecture details are secondary to performance" — layer count, head count, depth-width ratio matter very little within reasonable ranges; N, D, C (the three scale quantities) are the real drivers. This gave teams a clear directive: tuning architecture is less valuable than scaling. The paper also showed "performance improves as a power law with data, but with diminishing returns," and a transfer prediction (patterns on small models can be extrapolated to thousand-times larger models). Note: the experimental window was only ~1.5B parameters — extrapolation is a "belief," not a "proof" — Chinchilla is what corrected its underestimation on the data dimension.

6. Chinchilla: Training Compute-Optimal Large Language Models (2022, Hoffmann et al., DeepMind) ​

DimensionContent
Background & ProblemThe community, following Kaplan's conclusions, was blindly adding parameters (Gopher 280B, GPT-3 175B), but all were trained "data-starved." The paper challenged: under a fixed compute budget, what should the ratio of parameters to training tokens actually be?
Core MethodRedid the compute-optimal scan: fix total compute C, search for the optimal N and D combination. Conclusion: the optimal ratio is about 1:20 tokens per parameter (20 tokens per 1 parameter). Trained a 70B/1.4T token model (Chinchilla) per this ratio.
Key Experiments & NumbersChinchilla (70B, 1.4T tokens) outperformed Gopher (280B, 300B tokens) on average across 400+ evaluation tasks. Under the same compute, "smaller but trained longer" thoroughly beat "bigger but undertrained." Confirmed that a 70B model's loss can be comparable to a 540B model when data compensates.
Limitations① "Compute optimal" is loss-level optimal, not considering inference cost, deployment, or emergent abilities (small models aren't at a disadvantage for inference, but hard to say for capabilities); ② 1:20 is an empirical pattern — it drifts with different data quality; ③ Only answers "pretraining budget," not post-alignment training.
ExtensionsThe "add data or parameters first" practical conclusion in Scaling Laws; cost-effectiveness discussion on MoE in MoE Sparse Experts

Deep-dive takeaways: The 1:20 derivation is worth reading carefully: fix total compute C, model loss as a function of params N and data D, F(N, D), and find the (N, D) pair that minimizes loss. The conclusion was that the community was generally "undertrained" — flagship models like GPT-3 (175B/300B tokens, ~1:1.7) and Gopher (280B/300B, ~1:1) were far from data-sufficient. Per 1:20, they should have been trained on 1.4T+ tokens. This paper's practical impact was enormous: after this, all major model training planned data according to compute-optimal ratios — "data scarcity" became a more real bottleneck than "parameter scarcity." It's also the core basis for Pretraining: Data and Objectives.

4. Alignment and Instruction (1 Paper) ​

7. InstructGPT: Training Language Models to Follow Instructions with Human Feedback (2022, Ouyang et al., OpenAI) ​

DimensionContent
Background & ProblemGPT-3 could generate but wasn't "obedient": didn't follow instructions, wasn't honest, could be toxic. The goal shifted to alignment — making models accountable for "helpfulness, honesty, and safety."
Core MethodRLHF three steps: ① SFT — supervised fine-tuning with human-written instruction-response pairs (~13K pairs); ② RM — train a reward model to score human preference comparisons (~33K preference pairs); ③ PPO — optimize policy with the RM as reward and KL penalty to constrain deviation from the initialization.
Key Experiments & NumbersIn human evaluation, the 1.3B InstructGPT output was preferred over the 175B GPT-3 output (win rate ~70%+); real API evaluation from user feedback was similarly better; "alignment tax" was small: average only ~0.4% drop on traditional benchmarks like SQuAD (some tasks even improved).
Limitations① Relies on human annotation — expensive, limited scale; ② Model "people-pleases" rather than "is correct" (over-refusal, sycophancy); ③ PPO training is unstable and complex to implement — this directly motivated DPO; ④ Aligns only text output, not tool use or multimodal.
ExtensionsRLHF principles and DPO comparison in Alignment: RLHF and DPO; product story in ChatGPT and Conversational Models; alignment failure cases in Safety and Risks

Deep-dive takeaways: The core insight is that alignment "organizes" existing capability rather than "creating" it — GPT-3 at 175B already "knew" what to do, it just didn't know how to be asked. Engineering details worth scrutinizing: SFT used only ~13K pairs (data quality matters far more than quantity); RM used ~33K human preference pairs; PPO added KL penalty to prevent the policy from deviating too far from initialization. The "alignment tax" (slight drop on traditional benchmarks) was later shown to be mitigatable through better data and multi-task training. ChatGPT is the direct product of this technology plus a conversational product — the story is in ChatGPT and Conversational Models.

5. Applications and Systems (3 Papers) ​

8. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020, Lewis et al., Facebook AI) ​

DimensionContent
Background & ProblemPurely parametric models "memorize" knowledge in their weights: knowledge cutoff, expensive to update, factually error-prone. The paper asks: can the model retrieve from an external corpus and make knowledge external?
Core MethodRAG: a retriever (DPR dense retrieval) + generator (BART-large) trained jointly. Two decoding strategies: RAG-sequence (generate full sequence per document) and RAG-token (sample tokens at each step according to document distribution). Documents are concatenated with input as conditions; end-to-end trainable.
Key Experiments & NumbersOpen-domain QA: RAG-sequence surpassed purely parametric T5-11B on Natural Questions; set SOTA on FEVER (fact verification) and other tasks; joint training of retriever and generator brought additional gains.
Limitations① Retrieval quality is the ceiling — if it can't retrieve, it can't generate; ② Two-stage latency is higher than pure generation; ③ Limited benefit for "open-domain chit-chat" tasks that don't need knowledge; ④ The generator can still ignore retrieval results and hallucinate — RAG isn't a panacea for hallucination (see Hallucination).
ExtensionsParadigm evolution (Naive/Advanced/Modular RAG) in case-studies RAG; engineering deployment in RAG in Practice; comparison with long context and fine-tuning in Context and Long Context

Deep-dive takeaways: RAG moved knowledge storage from "inside parameters" to "outside parameters." Two design choices are worth reading carefully: the retriever uses DPR (dual-tower dense retrieval) instead of BM25 because it needs to be end-to-end differentiable; decoding uses RAG-sequence (generate a full passage per document) and RAG-token (sample at each step according to document distribution) — the former is more stable on open-domain QA, the latter is better for step-by-step citation. Its limitations foreshadowed a common challenge in all RAG systems: retrieval quality is the ceiling, and the generator may ignore retrieval results. RAG remains the primary solution for reducing hallucination (see Hallucination), though its form has evolved from naive to advanced/modular.

9. LoRA: Low-Rank Adaptation of Large Language Models (2021, Hu et al., Microsoft) ​

DimensionContent
Background & ProblemFull-parameter fine-tuning of a 175B model: each downstream task requires storing a full set of new weights (~350GB per task). GPU memory and storage are both infeasible. Can we "only learn the delta"?
Core MethodFreeze pretrained weights W₀, learn a low-rank delta ΔW = B·A (A, B are low-rank matrices, rank r ≪ d). Forward pass: W = W₀ + (α/r)·BA. The delta can be merged back into W₀ at inference, adding zero extra latency.
Key Experiments & NumbersOn GPT-3 175B, trainable parameters dropped from 175B to ~4.6M (a reduction of ~10,000×). On RoBERTa, GLUE results were comparable or better than full-parameter fine-tuning (e.g., surpassed on MRPC); rank r is insensitive within 1–64 (matrix decomposition works down to r=1).
Limitations① The low-rank assumption itself limits the representational space that can be learned (may not be enough for complex tasks); ② Hyperparameters (r, α, target_modules) require tuning; ③ The paper's scope is "adaptation," not full pretraining from scratch.
ExtensionsPrinciples and hyperparams in Fine-Tuning: SFT and PEFT; end-to-end practice in Fine-Tuning Practice: Full LoRA Workflow; QLoRA quantized version in Frontier Trends

Deep-dive takeaways: The low-rank assumption has experimental backing: performing SVD on updates to pretrained weights reveals that the dominant singular values are few — the delta does live in a low-rank subspace. A practical conclusion worth remembering: rank r between 1 and 64 makes almost no performance difference (experiments on GPT-3 in the paper), so "higher r is more stable" is a common misconception; α controls scaling, and in practice α/r is often kept constant. LoRA's significance isn't just saving GPU memory — it made "one base + multiple adapters" feasible, directly spawning the entire parameter-efficient fine-tuning family (QLoRA, DoRA, etc.) and forming the theoretical foundation for Fine-Tuning: SFT and PEFT.

10. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (2022, Dao et al., Stanford) ​

DimensionContent
Background & ProblemAttention O(n²) memory: at n=64k, a single fp32 attention matrix (64k × 64k) requires ~16GB. The bottleneck isn't just compute — it's memory bandwidth (HBM) — standard implementations write S and P back to HBM and read them back.
Core MethodI/O-aware exact attention: ① Tiling — partition Q, K, V into blocks and compute on SRAM; ② Recomputation — recompute the attention matrix during backpropagation, avoiding O(n²) intermediate storage; ③ Use an online update trick for exact softmax to maintain bit-accurate consistency with standard attention.
Key Experiments & NumbersMemory reduced from O(n²) to O(n). Compared to standard PyTorch attention, end-to-end training speedup: BERT-large ~15%, GPT-2 ~3×, Long-Range Arena ~2.4× (attention operator itself up to ~7.6×, increasing with sequence length).
Limitations① Only optimizes "standard attention" form; not directly applicable to sparse/linear attention; ② Engineering implementation depends on GPU architecture details; ③ Doesn't reduce attention compute complexity (still O(n²) compute).
ExtensionsAttention complexity and long-context engineering challenges in Context and Long Context; KV Cache optimization at inference time in Inference Fundamentals; FlashAttention-2 and PagedAttention in Frontier Trends

Deep-dive takeaways: The disruptive insight is "the bottleneck is I/O, not FLOPs": standard implementations write intermediate matrices S, P back to HBM and read them back, but HBM bandwidth is far slower than on-chip SRAM. Two technical pillars: tiling (block so SRAM can hold it) and recompute (recompute attention during backprop to save O(n²) intermediate storage). It maintains "exact" attention (numerically identical to standard implementation) — not approximate — so it can be drop-in-replaced into any attention. FA-2 further optimizes parallelization and tiling for another ~2× speedup. Reading this alongside understanding KV Cache connects the dots to Inference Fundamentals.

11. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022, Wei et al., Google) ​

DimensionContent
Background & ProblemLarge models perform poorly on math and multi-step reasoning: asking "what's the answer?" directly yields intuitive mistakes. The paper found: letting the model "think before answering" significantly elicits reasoning ability.
Core MethodChain-of-thought prompting: provide "intermediate reasoning steps" in few-shot examples (few-shot CoT); or simply append "Let's think step by step" (the prototype for zero-shot CoT). Reasoning steps aren't trained — they're elicited by prompting.
Key Experiments & NumbersOn PaLM 540B, GSM8K (elementary math) accuracy went from ~17.9% to ~56.9%; dramatic improvement on multi-subject exam benchmarks too. Scale threshold: models below ~100B get little to no CoT benefit — sometimes negative — classic evidence of "emergent abilities."
Limitations① Sensitive to model scale — ineffective for small models; ② CoT can be "confidently wrong" — reasoning process looks reasonable but conclusion is wrong; ③ Gains are most pronounced for tasks requiring rigor (e.g., math), limited for commonsense reasoning.
ExtensionsCoT and self-consistency engineering in Prompting and Prompting Practice; emergent ability debates in Scaling Laws

Deep-dive takeaways: "Emergence" is the most controversial and important observation in this paper: CoT shows no benefit (or even harm) below ~100B parameters, then suddenly improves significantly above it. This suggests reasoning ability isn't directly taught by the training objective — it's "revealed" once scale is sufficient. The way few-shot CoT examples are written matters — they should present a complete "step-by-step thinking" process. Later, zero-shot CoT ("Let's think step by step") was proven even simpler. Its successors: self-consistency uses multiple sampling + majority vote for further improvement; Tree of Thoughts (ToT) expands the search space. Note the boundary: CoT improves the "process" not "conclusion correctness" — cases where reasoning seems reasonable but the conclusion is wrong are common.

6. Cross-Paper Narrative: One Thread ​

These 11 papers together tell the complete LLM story — "decoder + scale + alignment":

text
Attention (2017) gave the "parallelizable, long-range" skeleton
   → GPT-1/2 (2018–2019) established the "decoder + pretraining" route
   → Scaling Laws (2020) gave the theoretical permission for "adding parameters"
   → GPT-3 empirical proof (2020) spread the scale belief (deep-dive in [Reading Paths](/papers/paths))
   → InstructGPT (2022) turned "can generate" into "can obey"
   → Chinchilla (2022) corrected the ratio: data and params are equally important
   → RAG/LoRA/CoT/FlashAttention (2020–2022) solved "affordable, usable, fast"

Three comparative relationships worth remembering:

ComparisonTakeaway
GPT (unidirectional decoder) vs BERT (bidirectional encoder)Route split between generation and understanding tasks; decoder ultimately won
Scaling Laws (add parameters) vs Chinchilla (add data)Under the same compute, params to tokens at ~1:20 is optimal
RLHF (InstructGPT, train reward model) vs Prompting (CoT, elicit reasoning)One changes weights, one changes context — they're stackable

The 11 papers don't need to be read sequentially by number. This grouping is more efficient (numbers correspond to sections on this page):

StageRecommended GroupReason
Step 11 (Attention) + 2 (GPT-1)Build the architecture skeleton first
Step 24 (BERT) + 3 (GPT-2)Two routes face each other like mirrors
Step 35 (Scaling) + 6 (Chinchilla)Scale thread: belief first, then correction
Step 47 (InstructGPT)The turning point from scale to alignment
Step 511 (CoT) + 8 (RAG)Applications: prompting and retrieval
Step 69 (LoRA) + 10 (FlashAttention)Systems: fine-tuning and acceleration

Within a group, read comparatively (e.g., 1 and 10: "the same attention, from architecture to system"). Between groups, order is not interchangeable: scale without architecture is a house of cards, alignment without scale has nothing to align.

8. Citation Relationships Between Papers ​

The value of deep-reading is seeing "who stands on whose shoulders":

PaperStands on the shoulders ofInspired
AttentionBahdanau attention, Seq2SeqAlmost all subsequent papers
GPT-1Attention (removed encoder)GPT-2/3, the entire decoder family
BERTAttention + ELMoRoBERTa, T5, the encoder family
Scaling LawsScale observations from GPT-2/3Chinchilla, all training budget planning
ChinchillaScaling LawsModern training ratio standards
InstructGPTRLHF (2017) + GPT-3ChatGPT, DPO, Constitutional AI
LoRAAdapter (2019)QLoRA, PEFT family
FlashAttentionAlgorithm + systems crossoverFA-2, PagedAttention
CoTGPT-3 few-shotSelf-consistency, ToT, o1
RAGDPR, BARTAdvanced RAG, multimodal retrieval

A starting point for following the vine

After completing this table, you have 10 "citation chains." From any paper, click into its citations, and you enter the expansion mode of the paper map — the map is alive.

9. Common Follow-Up Questions in Interviews ​

These 11 papers nearly cover all the high-frequency topics for ML interviews (full question bank in Interview Questions):

PaperPossible Interview Follow-UpKey Answer Points
AttentionWhy divide by √d_k? Why do multiple heads work?Prevent softmax saturation; heads attend to different subspaces
GPT-2What's the difference between zero-shot and few-shot?Prompt is the task description; few-shot adds examples
BERTIs NSP necessary?RoBERTa proved it can be removed; MLM is the real driver
Scaling LawsAdd data or add parameters?Check your current N:D ratio; reference 1:20
ChinchillaWhat is compute-optimal?Find optimal (N, D) pair under fixed compute
InstructGPTWhat does each of the RLHF three steps do?SFT → RM → PPO, KL penalty prevents drift
LoRAHow to set r and α?r between 1–64 makes almost no difference, keep α/r constant
FlashAttentionWhy does it save memory?I/O-aware: tiling + recomputation
CoTWhy doesn't it work for small models?Capability emergence has a scale threshold
RAGWhat if retrieval fails?Hybrid retrieval, query rewriting, reranking

10. What to Do After Deep-Reading ​

  1. Read the originals: This page is a navigation map — cross-reference each section with the original (links in the references below).
  2. Expand along citation chains: Continue from the "extensions" column on this page + each paper's Related Work.
  3. Take notes: Use the card note method to build a five-section card per paper: "problem / method / numbers / limitations / my scenario."
  4. Connect to the frontier: These are pre-2022 foundations. For 2023–2025 developments, see Frontier Trends.

Further Reading ​

References ​

Official originals for the 11 deep-dive papers (real arXiv links):