Appearance
Reading Paths
In one sentence: this page organizes paper reading around your goal, not around a bibliography — it answers "which papers should I read, in what order, how deeply, and how long will it take." It's for anyone with limited time who wants a roadmap rather than a book list. When you're done here, you can pick up the list and start reading right away, and you'll know when to stop and when to dig deeper.
The scariest part of reading papers isn't difficulty — it's the absence of a finish line. Readers without a route usually fall into one of two traps: they circle a survey forever, perpetually "still preparing," or they dive headfirst into the latest arXiv preprint and get shredded by the math. This page turns the three routes mentioned in Start Here into executable reading lists, with a "read to what depth" label and an estimated time for every paper.
The Three Routes at a Glance
text
Route 1: Two-Hour Quick Start (build "paper intuition")
DQN → PPO → AlphaGo (survey-style) → optional: one survey paper
Goal: hold your own in interviews, understand citations in blog posts
Route 2: Engineering Track (pick algorithms, reproduce results, diagnose)
6 classic close readings → RLHF (InstructGPT + DPO) → offline RL (CQL/IQL)
→ world models (DreamerV3/TD-MPC2) → reproduce one paper yourself
Goal: when a project problem hits, know which paper to look up
Route 3: Research Track (spot gaps, write related work)
Full paper map → reproduce 2–3 classics → follow new arXiv weekly
→ build a literature matrix by theme → find an attackable point
Goal: form your own research direction| Dimension | Route 1 | Route 2 | Route 3 |
|---|---|---|---|
| Goal | Build intuition | Support engineering decisions | Support research |
| Time | ~2 hours | 1–2 weeks (read while building) | 1–3 months |
| Paper count | 3–5 | 12–15 | 50+ |
| Reading depth | Abstract + method figure + conclusions | Method details + equations + reproduction | Everything + ablations + reproduction |
| Prerequisites | Basic Concepts | Route 1 + one project | Route 1 + Route 2 |
| Success check | Can answer "why" | Can reproduce a paper independently | Can write a page of related work |
These routes aren't mutually exclusive — they build on each other: Route 2 is Route 1 plus "engineering reproduction," and Route 3 is Route 2 plus "breadth + critique." At any point on the engineering or research tracks, you can drop back to Route 1 to refill your intuition.
Route 1: The Two-Hour Quick Start
This route covers only 3 classics + 1 optional survey, with the goal of building a feel for "what a paper looks like" within two hours. The order is deliberate: start with the simplest (DQN), move to the most practical (PPO), finish with the most spectacular (AlphaGo), and let a survey tie the three together at the end.
1. DQN: Mnih et al., Playing Atari with Deep Reinforcement Learning, 2013
- Link: arXiv:1312.5602 (the NIPS 2013 Deep Learning Workshop version, freely available; the 2015 Nature version is its engineering-hardened iteration)
- Read to what depth: abstract + method figure + intuition for the two engineering tricks
- What it's about: Using a deep network to approximate the Q-function and learning to play games directly from pixels. Two mechanisms you must take away — experience replay (store transitions in a buffer and sample them randomly, breaking the temporal correlation of samples) and target network (freeze an older Q-network to compute the targets, stabilizing training).
- Why it's first: It's the starting point of "deep RL," the plainest writing of the bunch, and its two tricks are the foundation of every deep RL algorithm that followed. Must-read sections: the two passages on experience replay and the target network. Skim the math (CNN architecture, loss function).
2. PPO: Schulman et al., Proximal Policy Optimization Algorithms, 2017
- Link: arXiv:1707.06347
- Read to what depth: Section 4 (the algorithm) + intuition for the clipped objective
- What it's about: A policy optimization algorithm that's "both stable and simple," limiting the step size of each update via clipping. The core formula is
L^CLIP(θ) = E_t[min(r_t(θ) Â_t, clip(r_t(θ), 1-ε, 1+ε) Â_t)], wherer_t(θ)=π_θ(a_t|s_t)/π_θold(a_t|s_t)is the probability ratio between the new and old policies. - Why it's second: It's the most-used algorithm today (including RLHF fine-tuning of language models) and one of the close readings in Classic Paper Readings. Must-read section: the passage on how the clipped objective approximates a "trust region" — why it can replace the complicated TRPO is the soul of this paper.
3. AlphaGo: Silver et al., Mastering the Game of Go with Deep Neural Networks and Tree Search, 2016
- Link: Nature article (no free arXiv version; AlphaZero has arXiv:1712.01815 as a substitute)
- Read to what depth: method figure + three-stage training (SL policy network → RL policy network → value network) + conclusions
- What it's about: Fusing supervised learning, reinforcement learning, and Monte Carlo tree search (MCTS) into a single system that defeats professional players. The three-stage training is the mechanism you must take away: first, supervise-train the policy network on 30 million human positions (predicting expert moves at ~57% accuracy); then have the policy network play against itself to improve (beating the supervised version at ~80% win rate); finally, train the value network on positions from the RL version's self-play games (predicting wins and losses, with error significantly lower than fast rollouts).
- Why it's last: It's the grandest of the three and information-dense — best read after the first two have built your algorithm intuition. Must-read sections: Figure 2 (system architecture) + the prose description of the three-stage training.
4. (Optional) Wrap-up survey: Li, Deep Reinforcement Learning: An Overview, 2017
- Link: arXiv:1701.07274
- Read to what depth: table of contents + first paragraph of each chapter
- Purpose: Ties the DQN, policy-gradient, and Actor-Critic lineages into one picture, and shows you "which other papers belong to these families."
Budgeting the two hours
DQN 40 minutes, PPO 50 minutes, AlphaGo 30 minutes, survey 20 minutes — about 140 minutes total. The "must-read sections" across all papers add up to fewer than 30 pages, so there's plenty of slack. Do not read every word — the discipline of Route 1 is "read only what needs reading."
Route 2: The Engineering Track
This track is for people who want to build things with RL. It uses the six close readings in Classic Paper Readings as its foundation, adds representative papers from four engineering-relevant directions, and ends by demanding that you reproduce one paper with your own hands. Total time: 1–2 weeks at 1–2 hours a day, verifying ideas in a project as you go.
1. Foundation: the 6 classic close readings (days 1–3)
The Classic Paper Readings page covers six papers in depth: Bellman 1957 (dynamic programming), Watkins 1989/1992 (Q-learning), Mnih 2015 (DQN), Silver 2016 (AlphaGo), Schulman 2017 (PPO), and Ouyang 2022 (InstructGPT). Spend the first three days of Route 2 reading them in order. The point isn't memorizing conclusions — it's understanding each algorithm as an answer to "what problem did it fix in the algorithm before it":
text
Bellman (MDP and the value equation, theoretical foundation)
└─> Q-learning (tabular, the off-policy revolution)
└─> DQN (tabular → neural network, fixes the state explosion)
└─> PPO (value-based → policy-based convergence, fixes stability)
└─> AlphaGo (algorithmic fusion + search, engineering tour de force)
└─> InstructGPT (RL migrates to language models)Each paper's "how to answer it in an interview" talking points are on that page — test yourself as you finish each one.
2. RLHF: InstructGPT and DPO (day 4)
- InstructGPT: Ouyang et al., Training Language Models to Follow Instructions with Human Feedback, 2022, arXiv:2203.02155. Three stages: SFT → reward model (human preference pairs) → PPO fine-tuning (with KL penalty). A must-read for LLM-related engineering; the RLHF concept page has a digest version.
- DPO: Rafailov et al., Direct Preference Optimization, 2023, arXiv:2305.18290. Optimize the policy directly on preference pairs, skipping both the reward model and PPO. "Simple to train" is a huge practical advantage, but its limitations are discussed in Frontier Progress.
3. Offline RL: pick CQL or IQL (day 5)
- CQL: Kumar et al., Conservative Q-Learning for Offline Reinforcement Learning, 2020, arXiv:2006.04779. Penalizes the value estimates of out-of-distribution (OOD) actions in the Q-learning objective.
- IQL: Kostrikov et al., Offline Reinforcement Learning with Implicit Q-Learning, 2022, arXiv:2110.06169. Uses expectile regression to learn only the support of the value function, never querying OOD actions.
- How to choose: If your project is "we have historical logs and want to learn a policy" (recommendation, risk control) → read IQL or CQL; if it's purely online training → skip this section. Background on the Offline RL concept page.
4. World models: pick DreamerV3 or TD-MPC2 (day 6)
- DreamerV3: Hafner et al., Mastering Diverse Domains through World Models, 2023, arXiv:2301.04104. One hyperparameter setting across 150+ tasks — the latent-space world-model philosophy plus fixed hyperparameters as an engineering doctrine.
- TD-MPC2: Hansen et al., Scalable, Robust World Models for Continuous Control, 2023, arXiv:2310.16828. A single multi-task world model for continuous control.
- How to choose: Vision-heavy or discrete tasks → DreamerV3; robotic continuous control → TD-MPC2. Background on the Model-Based RL concept page.
5. The mandatory step: reproduce one paper yourself (days 7–14)
Reproduction is the single most important success criterion of Route 2. Pick the paper closest to your work:
| Project context | Recommended reproduction |
|---|---|
| Games / Atari | DQN (adapt the skeleton from the Gymnasium Tutorial) |
| Continuous control / robotics | PPO or SAC (SB3 has reference implementations) |
| LLM alignment | DPO (the llm-alignment case study has a hands-on path) |
| Offline decision-making | IQL (run it on D4RL data) |
The discipline of reproduction is in Reading Discipline & FAQ: run the official implementation first to confirm you can reproduce, then write your own from scratch. Don't lock yourself in a room and reinvent the wheel on day one.
Route 3: The Research Track
This track is for people under paper pressure (grad school, publishing, hunting for a research direction). Its goal isn't to "finish reading" — it's to finish reading able to spot problems others have missed.
1. Phase 1: the full paper map (weeks 1–2)
Work through all six branches of the Paper Map end to end: dynamic programming, tabular methods, policy gradients, deep RL, RLHF, and multi-agent. For each branch, read the abstract + method figure of every key paper, and read the true classics in full. The goal isn't to memorize details — it's to build the coordinate system of "who has attacked this field from what angle." Without that coordinate system, you cannot see a gap.
2. Phase 2: reproduce 2–3 classics (weeks 3–4)
The close-reading + reproduction combo: pick 2–3 classic papers closest to your potential direction, run the official implementation first, then write one from scratch and run ablations (remove a trick and watch what happens). This produces two research assets: a reusable codebase, and first-hand intuition about "which mechanism actually does the work." The latter is your confidence when writing related work and designing ablation experiments.
3. Phase 3: weekly arXiv tracking (ongoing)
Build a fixed weekly rhythm: subscribe to arXiv by keyword (e.g., reinforcement learning, RLHF, world models) or check Papers with Code's Trending page. Each week, read abstracts of 3–5 new papers and one paper in full. Frontier Progress lists six frontier directions of the 2020s that can serve as your classification framework.
4. Phase 4: literature matrix and gap mining (from week 5)
Maintain a "literature matrix" as a table: rows are related papers, columns are (problem setting / method family / experimental environments / limitations / relation to my work). When the same limitation shows up in row after row (say, "every method is validated only on MuJoCo"), you've found an attackable point. This is the most underrated and most productive habit in research training.
text
Literature matrix example (excerpt):
┌────────────┬──────────────────┬────────────────┬──────────────────────┐
│ Paper │ Method family │ Limitation │ Relation to my work │
├────────────┼──────────────────┼────────────────┼──────────────────────┤
│ IQL │ Implicit Q │ D4RL only │ My setting is also │
│ │ │ │ offline │
│ CQL │ Conservative │ α needs tuning │ Can compare with IQL │
│ ... │ │ │ │
└────────────┴──────────────────┴────────────────┴──────────────────────┘Reading-Depth Levels for Every Paper
Before picking up any paper, decide which level you'll read it to — this avoids both extremes (one glance and done / grinding to the bitter end):
| Level | Covers | Time | When it's enough |
|---|---|---|---|
| L0 Title & abstract | Title + abstract + conclusions | 5 min | Deciding "should I read this at all" |
| L1 One-line positioning | L0 + method figure + contribution list | 15 min | Writing surveys; telling a colleague "what this paper does" |
| L2 Mechanism level | L1 + intuition for core equations + key experiments | 40–60 min | Interviews, choosing algorithms, understanding design rationale |
| L3 Detail level | L2 + derivations + all experiment tables + ablations | 2–4 hours | Reproduction, related work, critical reading |
| L4 Reproduction level | L3 + writing your own implementation and getting it to run | Days–weeks | Research, entering a new direction |
Route 1 reads everything at L2; Route 2 reads classics at L3 and frontier papers at L2; Route 3 reads relevant papers at L3–L4. Every paper on the Paper Map and in the Classic Readings page defaults to at least L2.
Master Time Budget
| Route | Total time | Reading | Reproduction | Success check |
|---|---|---|---|---|
| 1. Two-hour quick start | 2 hours | 3–5 papers × L2 | None | Can explain DQN's two tricks, the PPO clip intuition, and AlphaGo's three-stage training |
| 2. Engineering track | 1–2 weeks | 12–15 papers (3 days classics + 4 days frontier) | 1 full reproduction | Can locate the relevant paper when a project problem hits; can reproduce a paper independently |
| 3. Research track | 1–3 months | Full paper map + 3–5 papers weekly | 2–3 papers | Can write related work; can find at least one gap |
Common failure modes
Route 2's most common failure is "read a stack of papers, reproduced none" — reading 15 papers gives you zero engineering ability, while reproducing one gives you plenty. If you only have time for one thing, choose reproduction. Route 3's most common failure is the mirror image: "reproduce without reading the map" — redoing in a coordinate-free corner what someone already did, and wasting months discovering the answer was in the literature all along.
Further Reading
- Start Here — the section's front door: the positioning of the three routes and the "take away mechanisms, not numbers" principle.
- Classic Paper Readings — the foundation of Route 2: six milestone papers with close readings and interview talking points.
- Paper Map — the coordinate system for Route 3: one-line positioning of 40+ papers across six branches.
- Frontier Progress — the telescope for Routes 2/3: six frontier directions of the 2020s and the representative work in each.
- Reading Discipline & FAQ — the three-column note-taking method, reproduction discipline, and what to do when you can't understand something.
- Glossary — quick lookup for unfamiliar terms while reading.
References
- Mnih, V., et al. (2013). Playing Atari with Deep Reinforcement Learning. arXiv:1312.5602. https://arxiv.org/abs/1312.5602
- Schulman, J., et al. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347. https://arxiv.org/abs/1707.06347
- Silver, D., et al. (2016). Mastering the Game of Go with Deep Neural Networks and Tree Search. Nature 529:484–489. https://www.nature.com/articles/nature16961
- Silver, D., et al. (2017). Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm (AlphaZero). arXiv:1712.01815. https://arxiv.org/abs/1712.01815
- Li, Y. (2017). Deep Reinforcement Learning: An Overview. arXiv:1701.07274. https://arxiv.org/abs/1701.07274
- Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback (InstructGPT). arXiv:2203.02155. https://arxiv.org/abs/2203.02155
- Rafailov, R., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290. https://arxiv.org/abs/2305.18290
- Kumar, A., et al. (2020). Conservative Q-Learning for Offline Reinforcement Learning. arXiv:2006.04779. https://arxiv.org/abs/2006.04779
- Kostrikov, I., et al. (2022). Offline Reinforcement Learning with Implicit Q-Learning. arXiv:2110.06169. https://arxiv.org/abs/2110.06169
- Hafner, D., et al. (2023). Mastering Diverse Domains through World Models (DreamerV3). arXiv:2301.04104. https://arxiv.org/abs/2301.04104
- Hansen, N., et al. (2023). TD-MPC2: Scalable, Robust World Models for Continuous Control. arXiv:2310.16828. https://arxiv.org/abs/2310.16828