Appearance
Reading Paths
In one sentence: reading papers is not about "how many you've read" but about "taking the right path" — first pick your main line, then decide which papers to read and how deep to go. This page arranges the papers covered in this section into five thematic routes, each with a reading order and a self-check of what you should be able to answer when you're done.
1. Why Read Along Paths: An Ocean of Papers Needs Navigation
Start with a few numbers: the cs.LG (machine learning) category alone on arXiv gains more than 20,000 new papers a year; add in cs.CL (NLP), cs.CV (vision), and cs.AI, and hundreds of new papers go online every day. At this scale, two brutal facts hold:
- No one can read them all. Even top researchers close-read at most three to five papers a week, filtering the rest by title and abstract alone.
- Not filtering is a disaster. Paper quality is wildly uneven: there are paradigm-shifting foundational works, engineering-tuning reports, and padded filler. Read indiscriminately, and three months later you'll be holding a pile of fragments, unable to describe the skeleton of the field.
That is what a "path" is for. A path = a goal + an order + a check standard:
No path: skim a diffusion paper today → read an agent paper tomorrow → three months later:
you've seen everything and can explain none of it
With path: pick your main line first (LLMs? generation? engineering? alignment?)
→ read in the planned order (each paper builds on the previous one)
→ test yourself with the checks (can you answer those questions?)
→ three months later: a structured thematic map has taken shapeRead along a path, and no paper is an isolated file — each is a node on a map, where the question one paper raises is precisely the one the next paper answers. This interlocking of knowledge nodes makes your understanding compound. The site's companion pages support every step: The Paper Map provides the macro coordinate system, Classic Paper Deep-Dives the paragraph-by-paragraph microscope, and Frontier Developments the always-on radar.
2. The Five Routes at a Glance
Four main tracks follow today's hottest AI concepts (LLMs, generative AI, application engineering, alignment and safety), plus one "less but better" route as a fallback:
| Route | Sequence | Papers | Time budget | Best for |
|---|---|---|---|---|
| Route 1 The LLM Track | Transformer → BERT → GPT-3 → InstructGPT → Llama | 4 + 1 optional | 1 week skimmed / 2 weeks close-read | Everyone — understand "where language intelligence comes from" |
| Route 2 The Generative AI Track | GAN → DDPM → LDM → DiT / the Sora report | 3 + 2 optional | 1 week skimmed | Anyone interested in image/video generation |
| Route 3 The Application Engineering Track | RAG → ReAct → LoRA → GraphRAG | 3 + 1 optional | 4–5 days skimmed | Engineers and anyone putting AI into production |
| Route 4 The Alignment and Safety Track | InstructGPT → Constitutional AI → DPO | 3 | 3–4 days skimmed | Anyone who cares about trustworthy AI or is job-hunting for LLM roles |
| Route 5 Less but Better | Attention + GPT-3 + RAG + LoRA + DDPM | 5 | 1–2 hours per paper | Anyone short on time who wants one minimal, complete set |
Routes Stack — You Don't Have to Pick Just One
The five routes are not an either/or choice. Route 5 is the foundation for everyone — those five papers alone give you a feel for the whole landscape; Routes 1 and 3 are the engineer's standard kit; Route 4 is the main line for anyone going into alignment research. Suggested order: walk Route 5 first to build confidence, then enter the other tracks according to your goal.
3. Route 1: The LLM Track
Goal: understand the full causal chain of "why large language models work" — from the attention mechanism and the pretraining paradigm to scaling laws, alignment tuning, and open-source reproduction.
The papers and the logic of the sequence
| # | Paper (authors, year) | Contribution in one sentence | What problem from the previous paper it solved |
|---|---|---|---|
| 1 | Attention Is All You Need (Vaswani et al., 2017) | Replaced recurrence with self-attention and introduced the Transformer | Solves the RNN's inability to parallelize and its fading long-range dependencies |
| 2 | BERT (Devlin et al., 2018) | Bidirectional Transformer + masked language model (MLM) | Solves "how to pretrain on vast amounts of unlabeled text" |
| 3 | GPT-3 (Brown et al., 2020) | 175 billion parameters; demonstrated in-context learning | Solves "downstream tasks still need fine-tuning" — at scale, few-shot examples are enough |
| 4 | InstructGPT (Ouyang et al., 2022) | Used RLHF to make the model "understand people and speak like one" | Solves "powerful pretrained models that don't listen" — that is, alignment |
| 5 (optional) | Llama (Touvron et al., 2023) | Openly reproducible foundation models at 7B–65B scale | Solves "frontier models are closed off and can't be studied" — the starting point of the open-source ecosystem |
Suggested reading order
- Paper 1, Attention: read the abstract, Figure 1 (the encoder-decoder architecture), and the multi-head attention formulas in Section 3.2, alongside Jay Alammar's The Illustrated Transformer walkthrough (see References at the end). Grab one mechanism this week: why Q/K/V attention captures long-range dependencies in parallel.
- Paper 2, BERT: read the abstract, the two pretraining tasks (MLM and NSP), and the results tables. Hold on to one sentence: pretrain on massive text with "fill-in-the-blank" to get general-purpose representations, then fine-tune for the downstream task.
- Paper 3, GPT-3: read the abstract, the few-shot setup in Section 3, and the scaling curves. Accept one intuition: parameters × data × compute are the three dials of language intelligence.
- Paper 4, InstructGPT: read the abstract and the three-stage RLHF diagram in Section 3.1. Grab the sequence: supervised fine-tuning first, then a reward model, then reinforcement-learning optimization.
- Paper 5, Llama (optional): the abstract plus the training details (data mix, pretraining setup) are enough to learn how an open-source LLM gets built.
The Post-Reading Check
After these four (or five), you should be able to answer without looking at your notes: ① Why is the Transformer better suited to long text than RNNs? (parallelism + direct modeling of dependencies at any distance) ② What is the essential difference between BERT's and GPT's pretraining objectives? (bidirectional fill-in-the-blank vs. left-to-right next-word prediction) ③ What is in-context learning, and why was few-shot emergence not expected before GPT-3? (the definition of in-context learning and the effect of scale) ④ What is the order of RLHF's three stages, and what does each stage do? (SFT → reward model → PPO optimization)
If you can't answer them all, go back to the corresponding paper and reread that section; if you can answer all of them, the skeleton of this track has grown into your head.
Pair It with the Site's Resources
For the mechanics in depth, see Transformers and Attention and Large Language Models; for the product side, see ChatGPT and Conversational AI to understand how InstructGPT became a product.
4. Route 2: The Generative AI Track
Goal: understand "how machines create images and video from scratch" — from adversarial games to denoising diffusion, then into latent space and video generation.
The papers and the logic of the sequence
| # | Paper (authors, year) | Contribution in one sentence | What problem from the previous paper it solved |
|---|---|---|---|
| 1 | GAN (Goodfellow et al., 2014) | Generator and discriminator trained against each other to produce realistic images | Solves "generative models lacked a differentiable objective" — use the discriminator as the loss function |
| 2 | DDPM (Ho et al., 2020) | Add noise forward step by step, denoise backward — the diffusion model | Solves GAN's unstable training and mode collapse — a simple, stable objective |
| 3 | LDM / Stable Diffusion (Rombach et al., 2022) | Moves the diffusion process into latent space | Solves "pixel-space diffusion is too expensive" — a low-dimensional latent space is fast and good |
| 4 (optional) | DiT (Peebles et al., 2023) | Diffusion on a Transformer backbone (Diffusion Transformer) | Solves "U-Net's limited scalability" — the Transformer architecture improves smoothly with scale |
| 5 (optional) | The Sora technical report (OpenAI, 2024) | A video generation model framed as a "world simulator" | From generating static images to temporally consistent video generation |
Suggested reading order
- Paper 1, GAN: read the abstract and the game-theoretic intuition of the
min maxobjective. You only need the forger-vs-detective loop, and why it tends to mode collapse. - Paper 2, DDPM: read the abstract, Figure 2 (the noising/denoising illustration), and the training-algorithm pseudocode. Grab the intuition: forward, the image is noised bit by bit into pure noise; backward, the model learns to walk the noise back to an image, step by step. The training objective is just a simple denoising loss — far more stable than GAN.
- Paper 3, LDM: read the abstract and Figure 3 (the architecture). Grab the core idea: first use an autoencoder to compress the image into a low-dimensional latent space, then run diffusion there — the compute cost drops enormously, and consumer GPUs can play.
- Paper 4, DiT (optional): read the abstract and the scaling curves to see why "diffusion + Transformer" became the mainstream of video generation after 2024.
- Paper 5, the Sora report (optional): the body of the report is enough — it is a technical note rather than a rigorous paper; focus on the "world simulator" positioning and its own account of limitations.
The Post-Reading Check
After at least the first three, you should be able to answer: ① What is the core tension in GAN training? (the generator–discriminator min-max game; mode collapse) ② Why is diffusion's "noise forward, denoise backward" so stable to train? (every step's objective is denoising, so the gradient signal is clean) ③ Why does LDM move diffusion into latent space? (pixel space is too high-dimensional and computationally explosive; latent space is low-dimensional yet preserves semantics) ④ In one sentence: the rough pipeline behind a Stable Diffusion image? (encode the text → denoise in latent space → decode back to pixels)
If an answer won't come, go back to the figures in the corresponding paper; for an illustrated supplement to this track, see Diffusion Models and Generative AI.
5. Route 3: The Application Engineering Track
Goal: understand "how methods from papers become shippable product capabilities" — retrieval, agents, efficient fine-tuning; this is the set most often asked about when landing AI in production.
The papers and the logic of the sequence
| # | Paper (authors, year) | Contribution in one sentence | What problem from the previous paper it solved |
|---|---|---|---|
| 1 | RAG (Lewis et al., 2020) | Combines retrieval with generation so the model can cite external knowledge | Solves "a model's knowledge is frozen and it hallucinates" — bring in an external knowledge base |
| 2 | ReAct (Yao et al., 2022) | An alternating loop of reasoning (Thought) and acting (Action) | Solves "reasoning from parameters alone, with no contact with the outside world" — lets the model call tools |
| 3 | LoRA (Hu et al., 2021) | Efficient fine-tuning via low-rank decomposition, training only a small set of parameters | Solves "full fine-tuning is too expensive to afford or to store" — freeze the backbone and learn only the delta |
| 4 (optional) | GraphRAG (Microsoft, 2024) | Indexes with a knowledge graph before retrieving | Solves "vector retrieval can't answer global questions" — graph structure adds relational reasoning |
Suggested reading order
- Paper 1, RAG: read the abstract and the architecture diagram in Section 3.2. Grab the two pipelines, "retriever + generator", and understand why "retrieve first, then generate" reduces hallucination. For the engineering details (chunking, embeddings, vector stores), see Build a RAG App from Scratch.
- Paper 2, ReAct: read the abstract and the worked example in Figure 1. Grab the
Thought → Action → Observationloop — it is the theoretical prototype behind every engineering implementation of AI agents. - Paper 3, LoRA: read the abstract and the experiments in Section 4. Grab the core intuition: weight updates are usually low-rank, so they can be factored into two small matrices to be learned — trainable parameters shrink by a factor of 10,000 with no loss in score. To put it into practice, see Fine-Tune Your Own LLM.
- Paper 4, GraphRAG (optional): read the abstract and the community-detection / hierarchical-summarization pipeline, and understand why it beats naive RAG on "global questions" (like "what are these documents about overall?").
The Post-Reading Check
After the first three, you should be able to answer: ① What are RAG's two stages, and why does it ease hallucination? (retrieved evidence constrains generation) ② In ReAct's loop, what do reasoning and action each contribute? (Thought decides, Action calls tools, Observation feeds back) ③ Why does LoRA save GPU memory, and what is the intuition behind low-rank decomposition? (the delta matrix ≈ two small matrices, A × B) ④ Bonus question for engineers: how does poor retrieval quality drag down generation in RAG? (bad evidence misleads more than no evidence)
If you can answer all of them, you have the "paper → engineering solution" translation ability — exactly the core of what engineering interviews probe.
6. Route 4: The Alignment and Safety Track
Goal: understand "how to make models both capable and obedient" — from RLHF to self-alignment, then to DPO, which gets by without reinforcement learning. This is a high-frequency topic in LLM-role interviews and in AI Safety and Governance discussions.
The papers and the logic of the sequence
| # | Paper (authors, year) | Contribution in one sentence | What problem from the previous paper it solved |
|---|---|---|---|
| 1 | InstructGPT (Ouyang et al., 2022) | Three-stage RLHF: SFT → reward model → reinforcement learning | Solves "powerful pretrained models that don't follow instructions" |
| 2 | Constitutional AI (Bai et al., 2022) | A set of principles by which the model critiques and corrects itself | Solves "RLHF depends on massive human labeling" — rules replace the labelers |
| 3 | DPO (Rafailov et al., 2023) | Reduces RLHF to a preference-classification loss | Solves "reinforcement learning is engineering-heavy and unstable" — direct optimization without RL |
Suggested reading order
- Paper 1, InstructGPT: read the abstract and the pipeline in Section 3.1. Draw the three RLHF stages as a flowchart — it is the baseline for understanding all later alignment work.
- Paper 2, Constitutional AI: read the abstract and Figure 1 (the supervised stage + the RL stage). Grab the "critique → revision" loop: the model critiques its own responses against the constitutional principles, rewrites them, and is then reinforced with preference learning.
- Paper 3, DPO: read the abstract and the core formula in Section 3.2 (no need to grind through the derivation). Grab the conclusion: preference data can directly define the objective function, with no explicit reward model and no PPO — an order of magnitude simpler in engineering terms.
The Post-Reading Check
After all three, you should be able to answer: ① Why does RLHF need a reward-model step? (preferences are relative; you need an optimizable scalar objective) ② What does Constitutional AI replace with "principles", and at what cost? (it replaces human preference labeling; the cost is the quality of the principles themselves) ③ What is DPO's core simplification over RLHF? (it skips the explicit reward model and reinforcement learning, optimizing a classification loss directly) ④ What is the shared starting point of all three? (making outputs match human/rule-based preferences — they differ only in the path they take)
For the complete conceptual picture of alignment, see Alignment: RLHF and DPO; for how interview questions are framed, see the Interview Question Bank.
7. Route 5: Less but Better
Goal: when time is extremely scarce, build a minimal yet complete mental model of "AI's hot concepts" from just 5 papers. Each of the 5 maps onto a pillar: mechanism (Attention), scale (GPT-3), engineering (RAG and LoRA), generation (DDPM).
| # | Paper | Which pillar it represents | Skim time |
|---|---|---|---|
| 1 | Attention Is All You Need (2017) | The foundational mechanism of modern LLMs | 1–2 hours |
| 2 | GPT-3 (2020) | Scaling laws and in-context learning | 1–2 hours |
| 3 | RAG (2020) | Retrieval augmentation: the engineering answer to hallucination | 1 hour |
| 4 | LoRA (2021) | Efficient fine-tuning: adapting large models on a small budget | 1 hour |
| 5 | DDPM (2020) | Diffusion models: the technical foundation of generative AI | 1–2 hours |
Reading order and method
For each paper, read only the abstract → introduction → figures and tables → conclusion, and don't grind through the formulas. When you finish each one, write down in a single sentence of your own what problem it solved. Once all five are done, you can already answer the three trunk questions of today's hot AI topics: why large models work, how to make them more useful, and how images get generated.
The Post-Reading Check (Baseline Version)
After the five papers, you should be able to explain each to a non-technical friend in one plain sentence: Attention is "letting the model look at the whole context"; GPT-3 is "feed it enough and it can do a bit of everything"; RAG is "when it doesn't know, it looks it up first"; LoRA is "changing a little is enough"; DDPM is "walking the noise back to an image, step by step". If you can put it in plain words, you've truly understood it.
8. How to Budget Your Reading Time
The scarcest resource in paper reading is time. Set the budget first, then the depth — don't let perfectionism wreck your rhythm:
| Reading mode | Time per paper | What to read | What you get out of it |
|---|---|---|---|
| Skim | 1–2 hours | Abstract + introduction + figures + conclusion | Can state what problem it solves, the core method in one sentence, and how well it works |
| Close read | 4–6 hours | Full text + method + ablations + limitations + notes | Can explain the mechanism, the ablation evidence, and the boundaries of applicability — even judge whether it can be reused |
| Study | 8+ hours | Close read + derivations + reproduction | Can spot flaws in the paper and propose improvements (for researchers) |
Two pieces of practical advice:
- The 80/20 rule: about 80% of papers deserve only a skim (1–2 hours), 15% deserve a close read (4–6 hours), and 5% deserve full study. Close reading is an investment; skimming is an expense — use the expense to maintain breadth and the investment to buy depth.
- One hour a day beats a seven-hour weekend cram: understanding papers takes overnight fermentation; small, steady doses are far more solid than a single blitz. For methods and note-taking in detail, see Reading Discipline and FAQ.
The Most Common Time Trap
"Close-read every paper" is the most expensive mistake — it caps you at three papers a month, after which you quit over the slow progress. The right posture: skim first to quickly shortlist the candidates worth a close read, then put 4–6 hours into just a few of them. To judge whether a paper is worth a close read, look up where it sits in the genealogy on The Paper Map.
Further Reading
- Start Here — the section's main entrance: why read papers, the four-step method, and tool recommendations
- The Paper Map — the macro coordinate system for the five routes: origins, milestones, and surveys
- Classic Paper Deep-Dives — paragraph-by-paragraph breakdowns of the classic papers behind Routes 1, 2, and 3
- Frontier Developments — the extension lines of the routes: a systematic look at the newest breakthroughs of the 2020s
- Reading Discipline and FAQ — the three-pass method, note-taking, and what to do when you're stuck
- Learning Path Overview — from papers back to the whole site: the complete knowledge-learning route
- A Brief History of the AI Boom — putting the papers on a longer timeline: the historical arc of AI
- Glossary — look up any term that blocks you while reading
- Curated Resource List — a vetted collection of paper code reproductions, datasets, and tools
References
Everything below is a real, publicly available resource for going deeper on your own:
- The arXiv preprint library — the first-release home of the vast majority of AI papers
- Vaswani et al. Attention Is All You Need (NeurIPS 2017) — the original Transformer paper
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers (2018)
- Brown et al. Language Models are Few-Shot Learners (NeurIPS 2020) — GPT-3
- Ouyang et al. Training language models to follow instructions with human feedback (2022) — InstructGPT
- Touvron et al. LLaMA: Open and Efficient Foundation Language Models (2023)
- Goodfellow et al. Generative Adversarial Nets (NeurIPS 2014) — GAN
- Ho et al. Denoising Diffusion Probabilistic Models (NeurIPS 2020) — DDPM
- Rombach et al. High-Resolution Image Synthesis with Latent Diffusion Models (CVPR 2022) — LDM / Stable Diffusion
- Peebles et al. Scalable Diffusion Models with Transformers (ICCV 2023) — DiT
- Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020) — RAG
- Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models (ICLR 2023)
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models (ICLR 2022)
- Microsoft. From Local to Global: A Graph RAG Approach to Query-Focused Summarization (2024) — GraphRAG
- Bai et al. Constitutional AI: Harmlessness from AI Feedback (2022)
- Rafailov et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model (NeurIPS 2023) — DPO
- OpenAI. Video generation models as world simulators (the Sora technical report, 2024)
- Jay Alammar. The Illustrated Transformer (2018) — the classic example of an illustrated Transformer close read
- The Annotated Transformer (Harvard NLP) — a close reading of the Transformer with line-by-line annotations and runnable code