Skip to content

Reading Paths

At a glance The ocean of papers is vast and deep, and diving in blind will only drive you away — this page lays out five thematic reading routes (LLMs, generative AI, application engineering, alignment and safety, plus a less-but-better fallback), each with a reading order, a self-check for when you finish, and a suggested time budget.

Reading Paths ​

In one sentence: reading papers is not about "how many you've read" but about "taking the right path" — first pick your main line, then decide which papers to read and how deep to go. This page arranges the papers covered in this section into five thematic routes, each with a reading order and a self-check of what you should be able to answer when you're done.

1. Why Read Along Paths: An Ocean of Papers Needs Navigation ​

Start with a few numbers: the cs.LG (machine learning) category alone on arXiv gains more than 20,000 new papers a year; add in cs.CL (NLP), cs.CV (vision), and cs.AI, and hundreds of new papers go online every day. At this scale, two brutal facts hold:

  1. No one can read them all. Even top researchers close-read at most three to five papers a week, filtering the rest by title and abstract alone.
  2. Not filtering is a disaster. Paper quality is wildly uneven: there are paradigm-shifting foundational works, engineering-tuning reports, and padded filler. Read indiscriminately, and three months later you'll be holding a pile of fragments, unable to describe the skeleton of the field.

That is what a "path" is for. A path = a goal + an order + a check standard:

No path:    skim a diffusion paper today → read an agent paper tomorrow → three months later:
            you've seen everything and can explain none of it

With path:  pick your main line first (LLMs? generation? engineering? alignment?)
            → read in the planned order (each paper builds on the previous one)
            → test yourself with the checks (can you answer those questions?)
            → three months later: a structured thematic map has taken shape

Read along a path, and no paper is an isolated file — each is a node on a map, where the question one paper raises is precisely the one the next paper answers. This interlocking of knowledge nodes makes your understanding compound. The site's companion pages support every step: The Paper Map provides the macro coordinate system, Classic Paper Deep-Dives the paragraph-by-paragraph microscope, and Frontier Developments the always-on radar.

2. The Five Routes at a Glance ​

Four main tracks follow today's hottest AI concepts (LLMs, generative AI, application engineering, alignment and safety), plus one "less but better" route as a fallback:

RouteSequencePapersTime budgetBest for
Route 1 The LLM TrackTransformer → BERT → GPT-3 → InstructGPT → Llama4 + 1 optional1 week skimmed / 2 weeks close-readEveryone — understand "where language intelligence comes from"
Route 2 The Generative AI TrackGAN → DDPM → LDM → DiT / the Sora report3 + 2 optional1 week skimmedAnyone interested in image/video generation
Route 3 The Application Engineering TrackRAG → ReAct → LoRA → GraphRAG3 + 1 optional4–5 days skimmedEngineers and anyone putting AI into production
Route 4 The Alignment and Safety TrackInstructGPT → Constitutional AI → DPO33–4 days skimmedAnyone who cares about trustworthy AI or is job-hunting for LLM roles
Route 5 Less but BetterAttention + GPT-3 + RAG + LoRA + DDPM51–2 hours per paperAnyone short on time who wants one minimal, complete set

Routes Stack — You Don't Have to Pick Just One

The five routes are not an either/or choice. Route 5 is the foundation for everyone — those five papers alone give you a feel for the whole landscape; Routes 1 and 3 are the engineer's standard kit; Route 4 is the main line for anyone going into alignment research. Suggested order: walk Route 5 first to build confidence, then enter the other tracks according to your goal.

3. Route 1: The LLM Track ​

Goal: understand the full causal chain of "why large language models work" — from the attention mechanism and the pretraining paradigm to scaling laws, alignment tuning, and open-source reproduction.

The papers and the logic of the sequence ​

#Paper (authors, year)Contribution in one sentenceWhat problem from the previous paper it solved
1Attention Is All You Need (Vaswani et al., 2017)Replaced recurrence with self-attention and introduced the TransformerSolves the RNN's inability to parallelize and its fading long-range dependencies
2BERT (Devlin et al., 2018)Bidirectional Transformer + masked language model (MLM)Solves "how to pretrain on vast amounts of unlabeled text"
3GPT-3 (Brown et al., 2020)175 billion parameters; demonstrated in-context learningSolves "downstream tasks still need fine-tuning" — at scale, few-shot examples are enough
4InstructGPT (Ouyang et al., 2022)Used RLHF to make the model "understand people and speak like one"Solves "powerful pretrained models that don't listen" — that is, alignment
5 (optional)Llama (Touvron et al., 2023)Openly reproducible foundation models at 7B–65B scaleSolves "frontier models are closed off and can't be studied" — the starting point of the open-source ecosystem

Suggested reading order ​

  • Paper 1, Attention: read the abstract, Figure 1 (the encoder-decoder architecture), and the multi-head attention formulas in Section 3.2, alongside Jay Alammar's The Illustrated Transformer walkthrough (see References at the end). Grab one mechanism this week: why Q/K/V attention captures long-range dependencies in parallel.
  • Paper 2, BERT: read the abstract, the two pretraining tasks (MLM and NSP), and the results tables. Hold on to one sentence: pretrain on massive text with "fill-in-the-blank" to get general-purpose representations, then fine-tune for the downstream task.
  • Paper 3, GPT-3: read the abstract, the few-shot setup in Section 3, and the scaling curves. Accept one intuition: parameters × data × compute are the three dials of language intelligence.
  • Paper 4, InstructGPT: read the abstract and the three-stage RLHF diagram in Section 3.1. Grab the sequence: supervised fine-tuning first, then a reward model, then reinforcement-learning optimization.
  • Paper 5, Llama (optional): the abstract plus the training details (data mix, pretraining setup) are enough to learn how an open-source LLM gets built.

The Post-Reading Check

After these four (or five), you should be able to answer without looking at your notes: ① Why is the Transformer better suited to long text than RNNs? (parallelism + direct modeling of dependencies at any distance) ② What is the essential difference between BERT's and GPT's pretraining objectives? (bidirectional fill-in-the-blank vs. left-to-right next-word prediction) ③ What is in-context learning, and why was few-shot emergence not expected before GPT-3? (the definition of in-context learning and the effect of scale) ④ What is the order of RLHF's three stages, and what does each stage do? (SFT → reward model → PPO optimization)

If you can't answer them all, go back to the corresponding paper and reread that section; if you can answer all of them, the skeleton of this track has grown into your head.

Pair It with the Site's Resources

For the mechanics in depth, see Transformers and Attention and Large Language Models; for the product side, see ChatGPT and Conversational AI to understand how InstructGPT became a product.

4. Route 2: The Generative AI Track ​

Goal: understand "how machines create images and video from scratch" — from adversarial games to denoising diffusion, then into latent space and video generation.

The papers and the logic of the sequence ​

#Paper (authors, year)Contribution in one sentenceWhat problem from the previous paper it solved
1GAN (Goodfellow et al., 2014)Generator and discriminator trained against each other to produce realistic imagesSolves "generative models lacked a differentiable objective" — use the discriminator as the loss function
2DDPM (Ho et al., 2020)Add noise forward step by step, denoise backward — the diffusion modelSolves GAN's unstable training and mode collapse — a simple, stable objective
3LDM / Stable Diffusion (Rombach et al., 2022)Moves the diffusion process into latent spaceSolves "pixel-space diffusion is too expensive" — a low-dimensional latent space is fast and good
4 (optional)DiT (Peebles et al., 2023)Diffusion on a Transformer backbone (Diffusion Transformer)Solves "U-Net's limited scalability" — the Transformer architecture improves smoothly with scale
5 (optional)The Sora technical report (OpenAI, 2024)A video generation model framed as a "world simulator"From generating static images to temporally consistent video generation

Suggested reading order ​

  • Paper 1, GAN: read the abstract and the game-theoretic intuition of the min max objective. You only need the forger-vs-detective loop, and why it tends to mode collapse.
  • Paper 2, DDPM: read the abstract, Figure 2 (the noising/denoising illustration), and the training-algorithm pseudocode. Grab the intuition: forward, the image is noised bit by bit into pure noise; backward, the model learns to walk the noise back to an image, step by step. The training objective is just a simple denoising loss — far more stable than GAN.
  • Paper 3, LDM: read the abstract and Figure 3 (the architecture). Grab the core idea: first use an autoencoder to compress the image into a low-dimensional latent space, then run diffusion there — the compute cost drops enormously, and consumer GPUs can play.
  • Paper 4, DiT (optional): read the abstract and the scaling curves to see why "diffusion + Transformer" became the mainstream of video generation after 2024.
  • Paper 5, the Sora report (optional): the body of the report is enough — it is a technical note rather than a rigorous paper; focus on the "world simulator" positioning and its own account of limitations.

The Post-Reading Check

After at least the first three, you should be able to answer: ① What is the core tension in GAN training? (the generator–discriminator min-max game; mode collapse) ② Why is diffusion's "noise forward, denoise backward" so stable to train? (every step's objective is denoising, so the gradient signal is clean) ③ Why does LDM move diffusion into latent space? (pixel space is too high-dimensional and computationally explosive; latent space is low-dimensional yet preserves semantics) ④ In one sentence: the rough pipeline behind a Stable Diffusion image? (encode the text → denoise in latent space → decode back to pixels)

If an answer won't come, go back to the figures in the corresponding paper; for an illustrated supplement to this track, see Diffusion Models and Generative AI.

5. Route 3: The Application Engineering Track ​

Goal: understand "how methods from papers become shippable product capabilities" — retrieval, agents, efficient fine-tuning; this is the set most often asked about when landing AI in production.

The papers and the logic of the sequence ​

#Paper (authors, year)Contribution in one sentenceWhat problem from the previous paper it solved
1RAG (Lewis et al., 2020)Combines retrieval with generation so the model can cite external knowledgeSolves "a model's knowledge is frozen and it hallucinates" — bring in an external knowledge base
2ReAct (Yao et al., 2022)An alternating loop of reasoning (Thought) and acting (Action)Solves "reasoning from parameters alone, with no contact with the outside world" — lets the model call tools
3LoRA (Hu et al., 2021)Efficient fine-tuning via low-rank decomposition, training only a small set of parametersSolves "full fine-tuning is too expensive to afford or to store" — freeze the backbone and learn only the delta
4 (optional)GraphRAG (Microsoft, 2024)Indexes with a knowledge graph before retrievingSolves "vector retrieval can't answer global questions" — graph structure adds relational reasoning

Suggested reading order ​

  • Paper 1, RAG: read the abstract and the architecture diagram in Section 3.2. Grab the two pipelines, "retriever + generator", and understand why "retrieve first, then generate" reduces hallucination. For the engineering details (chunking, embeddings, vector stores), see Build a RAG App from Scratch.
  • Paper 2, ReAct: read the abstract and the worked example in Figure 1. Grab the Thought → Action → Observation loop — it is the theoretical prototype behind every engineering implementation of AI agents.
  • Paper 3, LoRA: read the abstract and the experiments in Section 4. Grab the core intuition: weight updates are usually low-rank, so they can be factored into two small matrices to be learned — trainable parameters shrink by a factor of 10,000 with no loss in score. To put it into practice, see Fine-Tune Your Own LLM.
  • Paper 4, GraphRAG (optional): read the abstract and the community-detection / hierarchical-summarization pipeline, and understand why it beats naive RAG on "global questions" (like "what are these documents about overall?").

The Post-Reading Check

After the first three, you should be able to answer: ① What are RAG's two stages, and why does it ease hallucination? (retrieved evidence constrains generation) ② In ReAct's loop, what do reasoning and action each contribute? (Thought decides, Action calls tools, Observation feeds back) ③ Why does LoRA save GPU memory, and what is the intuition behind low-rank decomposition? (the delta matrix ≈ two small matrices, A × B) ④ Bonus question for engineers: how does poor retrieval quality drag down generation in RAG? (bad evidence misleads more than no evidence)

If you can answer all of them, you have the "paper → engineering solution" translation ability — exactly the core of what engineering interviews probe.

6. Route 4: The Alignment and Safety Track ​

Goal: understand "how to make models both capable and obedient" — from RLHF to self-alignment, then to DPO, which gets by without reinforcement learning. This is a high-frequency topic in LLM-role interviews and in AI Safety and Governance discussions.

The papers and the logic of the sequence ​

#Paper (authors, year)Contribution in one sentenceWhat problem from the previous paper it solved
1InstructGPT (Ouyang et al., 2022)Three-stage RLHF: SFT → reward model → reinforcement learningSolves "powerful pretrained models that don't follow instructions"
2Constitutional AI (Bai et al., 2022)A set of principles by which the model critiques and corrects itselfSolves "RLHF depends on massive human labeling" — rules replace the labelers
3DPO (Rafailov et al., 2023)Reduces RLHF to a preference-classification lossSolves "reinforcement learning is engineering-heavy and unstable" — direct optimization without RL

Suggested reading order ​

  • Paper 1, InstructGPT: read the abstract and the pipeline in Section 3.1. Draw the three RLHF stages as a flowchart — it is the baseline for understanding all later alignment work.
  • Paper 2, Constitutional AI: read the abstract and Figure 1 (the supervised stage + the RL stage). Grab the "critique → revision" loop: the model critiques its own responses against the constitutional principles, rewrites them, and is then reinforced with preference learning.
  • Paper 3, DPO: read the abstract and the core formula in Section 3.2 (no need to grind through the derivation). Grab the conclusion: preference data can directly define the objective function, with no explicit reward model and no PPO — an order of magnitude simpler in engineering terms.

The Post-Reading Check

After all three, you should be able to answer: ① Why does RLHF need a reward-model step? (preferences are relative; you need an optimizable scalar objective) ② What does Constitutional AI replace with "principles", and at what cost? (it replaces human preference labeling; the cost is the quality of the principles themselves) ③ What is DPO's core simplification over RLHF? (it skips the explicit reward model and reinforcement learning, optimizing a classification loss directly) ④ What is the shared starting point of all three? (making outputs match human/rule-based preferences — they differ only in the path they take)

For the complete conceptual picture of alignment, see Alignment: RLHF and DPO; for how interview questions are framed, see the Interview Question Bank.

7. Route 5: Less but Better ​

Goal: when time is extremely scarce, build a minimal yet complete mental model of "AI's hot concepts" from just 5 papers. Each of the 5 maps onto a pillar: mechanism (Attention), scale (GPT-3), engineering (RAG and LoRA), generation (DDPM).

#PaperWhich pillar it representsSkim time
1Attention Is All You Need (2017)The foundational mechanism of modern LLMs1–2 hours
2GPT-3 (2020)Scaling laws and in-context learning1–2 hours
3RAG (2020)Retrieval augmentation: the engineering answer to hallucination1 hour
4LoRA (2021)Efficient fine-tuning: adapting large models on a small budget1 hour
5DDPM (2020)Diffusion models: the technical foundation of generative AI1–2 hours

Reading order and method ​

For each paper, read only the abstract → introduction → figures and tables → conclusion, and don't grind through the formulas. When you finish each one, write down in a single sentence of your own what problem it solved. Once all five are done, you can already answer the three trunk questions of today's hot AI topics: why large models work, how to make them more useful, and how images get generated.

The Post-Reading Check (Baseline Version)

After the five papers, you should be able to explain each to a non-technical friend in one plain sentence: Attention is "letting the model look at the whole context"; GPT-3 is "feed it enough and it can do a bit of everything"; RAG is "when it doesn't know, it looks it up first"; LoRA is "changing a little is enough"; DDPM is "walking the noise back to an image, step by step". If you can put it in plain words, you've truly understood it.

8. How to Budget Your Reading Time ​

The scarcest resource in paper reading is time. Set the budget first, then the depth — don't let perfectionism wreck your rhythm:

Reading modeTime per paperWhat to readWhat you get out of it
Skim1–2 hoursAbstract + introduction + figures + conclusionCan state what problem it solves, the core method in one sentence, and how well it works
Close read4–6 hoursFull text + method + ablations + limitations + notesCan explain the mechanism, the ablation evidence, and the boundaries of applicability — even judge whether it can be reused
Study8+ hoursClose read + derivations + reproductionCan spot flaws in the paper and propose improvements (for researchers)

Two pieces of practical advice:

  1. The 80/20 rule: about 80% of papers deserve only a skim (1–2 hours), 15% deserve a close read (4–6 hours), and 5% deserve full study. Close reading is an investment; skimming is an expense — use the expense to maintain breadth and the investment to buy depth.
  2. One hour a day beats a seven-hour weekend cram: understanding papers takes overnight fermentation; small, steady doses are far more solid than a single blitz. For methods and note-taking in detail, see Reading Discipline and FAQ.

The Most Common Time Trap

"Close-read every paper" is the most expensive mistake — it caps you at three papers a month, after which you quit over the slow progress. The right posture: skim first to quickly shortlist the candidates worth a close read, then put 4–6 hours into just a few of them. To judge whether a paper is worth a close read, look up where it sits in the genealogy on The Paper Map.

Further Reading ​

  • Start Here — the section's main entrance: why read papers, the four-step method, and tool recommendations
  • The Paper Map — the macro coordinate system for the five routes: origins, milestones, and surveys
  • Classic Paper Deep-Dives — paragraph-by-paragraph breakdowns of the classic papers behind Routes 1, 2, and 3
  • Frontier Developments — the extension lines of the routes: a systematic look at the newest breakthroughs of the 2020s
  • Reading Discipline and FAQ — the three-pass method, note-taking, and what to do when you're stuck
  • Learning Path Overview — from papers back to the whole site: the complete knowledge-learning route
  • A Brief History of the AI Boom — putting the papers on a longer timeline: the historical arc of AI
  • Glossary — look up any term that blocks you while reading
  • Curated Resource List — a vetted collection of paper code reproductions, datasets, and tools

References ​

Everything below is a real, publicly available resource for going deeper on your own: