Skip to content

Start Here

On this page Why RL learners should read papers — understand principles instead of memorizing APIs, keep up with the frontier instead of chasing hype, and ace interviews; the three reading routes; a map of the section's five pages; "take away mechanisms, not numbers."

Start Here ​

In one sentence: this page is the front door of the papers/ (Paper Reading) section — it answers "do I really need to read papers to learn RL, which ones, in what order, and to what depth." It's for anyone who has just finished Basic Concepts and wants to level up, as well as people preparing for interviews or working on projects who need to follow the literature. When you're done here, you'll be able to pick a reading route that fits you — and know which page to open first.

For many people, learning RL goes like this: step one, set up the environment and run Gymnasium; step two, follow a blog post and copy some code. This step is necessary — but it comes with a hidden cost: you learn "how to get someone else's algorithm running," not "why the algorithm looks the way it does." A blog will tell you how to write PPO's clipped objective; only a paper will tell you the story of the "trust region" idea behind the clip, handed down from TRPO. This documentation's four content pages — Reading Paths, Paper Map, Classic Paper Readings, and Frontier Progress — plus this page's companion FAQ (Reading Discipline & FAQ), turn paper reading from "intimidating" into "you have a route, a map, and a method."

The Section's Mission: Three Reasons to Read Papers ​

Reason 1: Understand principles, not memorize APIs ​

Framework APIs change: the MlpPolicy in SB3 today may be rewritten for the JAX ecosystem tomorrow. But the Q-learning update, PPO's clipped objective, and AlphaGo's three-stage training haven't changed in twenty years. These "mechanisms" are the knowledge that actually transfers. Consider the table below:

How you learnWhat you gainShelf lifeCost
Running framework examplesAbility to reproduce an algorithmSix months to a year (until the framework updates)Low
Reading surveys / blog postsA bird's-eye impression of algorithm familiesOne to three yearsLow
Close-reading classic papersMechanisms, motivations, trade-offs, design rationaleFive years or more (practically never expires)High
Close-reading + reproductionAll of the above + the ability to modify, diagnose, and innovateA lifetimeVery high

Pay attention to the "shelf life" column. Before 2022, "got a DQN running" was still a resume highlight; in 2026, that line no longer belongs on a resume, because one model.learn() call in a framework does it for you. But the question "why does DQN need experience replay, and why does it need a target network" will forever separate people who memorized APIs from people who understand principles — it remains one of interviewers' favorite follow-up questions. See Classic Paper Readings.

Reason 2: Keep up with the frontier, not chase hype ​

Hype is news; the frontier is papers. The difference: hype has typically been second-hand-retold, simplified, or even distorted (e.g., "RLHF is just PPO with a reward model" — dropping the KL penalty, the very mechanism that prevents the policy from drifting off course). The frontier is the first-hand statement, complete with the authors' own caveats and failure reports.

text
How hype propagates:
paper → summary/blog retelling → tech-press coverage → secondhand chatter
      ↑                              ↑
  information decays: caveats   almost nothing left
  dropped, failures omitted,    but the headline and the
  numbers inflated              conclusion — even "when it
                                holds" is gone

How frontier reading works:
paper → abstract → introduction (motivation) → method → experiments → discussion
                                    ↑
  the further you go, the closer to the truth: the method's assumptions,
  the experiments' limitations, the authors' own assessment

So the right way to "keep up with the frontier" isn't scrolling tech media — it's reading a few representative arXiv papers each quarter (Frontier Progress has picked them for you). Take 2025's DeepSeek-R1: the media tells you "it used reinforcement learning to make a model reason," but the paper tells you the key mechanism was GRPO (Group Relative Policy Optimization) — no critic network; instead, advantages are normalized by within-group comparisons, which simplifies the training infrastructure for large-scale RL enormously. That mechanism is worth far more than the conclusion "R1 is impressive."

Reason 3: A hard advantage in interviews ​

RL interviews have a brutal reality: too many people memorized concepts; too few have read the originals. In the Interview Question Bank, questions like "why is PPO easier to implement than TRPO," "why does Q-learning overestimate values," and "how was AlphaGo's value network trained" can be answered at some level by nearly every candidate — but only those who have read the papers can supply the crucial details:

  • TRPO needs second derivatives and conjugate gradients, and in engineering terms requires maintaining a whole constrained-optimization stack; PPO approximates the trust region with a first-order clip — a few lines of code, easy to tune, easy to parallelize. That motivation is stated explicitly by the authors in the PPO paper.
  • Q-learning's overestimation comes from the max operator in its update: the same noise gets max-ed, and the expectation is systematically inflated; Double DQN breaks this by "letting the online network pick the action and the target network score it."
  • AlphaGo's value network was not trained on real game records — it was trained by regression on win/loss from positions in the RL policy network's own self-play games. That's the stage of the "three-stage training" people most often get wrong.

These details can't be retained from tutorials alone, because they "aren't the tutorial's focus — but they are precisely the paper's focus." See Classic Paper Readings.

The Three Reading Routes ​

People with different goals need completely different routes. The Reading Paths page gives the full lists; here's the overview:

RouteBest forTimeCore materialWhat you can do afterwards
Two-hour quick startTime-pressed; just want "paper intuition"~2 hoursAbstracts + method figures + conclusions of 3–5 classicsHold your own in interviews; understand citations in blog posts
Engineering trackBuilding RL into a project/product1–2 weeks (read while building)Classic readings + engineering-relevant frontier + reproductionChoose algorithms, reproduce results, diagnose training issues
Research trackWant to publish / work on a grad-school topic1–3 monthsFull paper map + latest arXiv + reproduction experimentsSpot gaps, write related work, design ablations
text
How the three routes relate (not exclusive — progressive):
Two-hour quick start ──┐
                       ├──> Engineering track ──> Research track
   only need to "understand"   need to "use it"     need to "see the problems"

My advice

If you're preparing for interviews, finish at least the "two-hour quick start" route, then pick the one paper in Classic Paper Readings most relevant to the role you're targeting (DQN family for recommendation roles, InstructGPT for LLM roles, PPO/SAC for robotics roles), and read it closely while reproducing the key equations.

A Map of the Section's Pages ​

The papers/ section has 6 pages; this page is the entrance, and the other 5 each have a job to do:

PageRoleQuestion it answersWhen to read it
Reading PathsRoadmapWhich papers should I read, in what order?Before you start reading papers
Paper MapPanoramaHow do seventy years of key RL papers line up by theme/time?When you want a global coordinate system
Classic Paper ReadingsClose readingHow to read each of the 6 papers that changed RL, and how to answer interview questions about themWhen preparing for interviews / building foundations
Frontier ProgressTelescopeWhat are the frontier directions of the 2020s, and which are real breakthroughs?When chasing hot topics or hunting for a research direction
Reading Discipline & FAQMethodologyWhat to do when the math doesn't parse? How to take notes? Which papers can I skip?When reading your first paper, and every time you get stuck
text
Typical flow:
                     ┌──────────────────────┐
   Start ──> Reading Paths ──> Classic Paper Readings / Paper Map ──> Frontier Progress
                  │            │                │
                  └────> Stuck? ──> Reading Discipline & FAQ
                                     │
                                     └──> Unknown term? ──> Glossary /resources/glossary

One note: terminology is the first hurdle of paper reading. When you hit an unfamiliar term, go straight to the Glossary (60+ entries, each a one-line definition + related pages) — much faster than a search engine.

The Single Most Important Piece of Advice: Take Away Mechanisms, Not Numbers ​

If you remember only one sentence from this section, make it this one: when you read a paper, take away the mechanism, not the number.

Numbers in papers come in two kinds:

  • Conclusion numbers: e.g., "AlphaGo beat Lee Sedol 4:1," "the 1.3B InstructGPT was preferred to the 175B GPT-3 by human evaluators." These are context-dependent — change the compute, the data, the evaluation protocol, and the numbers no longer hold. Getting the year or the score slightly wrong in an interview is really not fatal; not knowing why AlphaGo won is the fatal part.
  • Mechanistic knowledge: e.g., "Q-learning is off-policy because its update uses max_a' Q(s',a') regardless of how the behavior policy chooses actions," or "PPO's clip constrains the ratio r_t(θ) between new and old policies, not the distance between parameter vectors." This kind of knowledge doesn't depend on any specific number, and it never expires.
text
The same paper, two ways to read it:

Reading A (remember numbers): AlphaGo beat Lee Sedol 4:1, using 1,202 CPUs
    and 176 GPUs.
    → Asked "why could AlphaGo win?" in an interview, you can only say
      "it was really strong."

Reading B (remember mechanisms): AlphaGo = three-stage training (SL on human
    games → RL self-play reinforcement → value network evaluating positions)
    + MCTS at inference time combining "policy prior × value estimate."
    → Asked "why could AlphaGo win?", you can lay out the full logic of
      "trading learned knowledge for compute," and point out that AlphaZero
      later removed the human game records altogether.

One more example, closer to everyday work. The DeepSeek-R1 paper contains a detail many people love to chat about: during training, the model "spontaneously" learned to reflect for longer, which the authors called the "aha moment." That detail makes good small talk, but the actually transferable mechanism is this: R1's RL stage didn't use standard PPO — it used GRPO, generating a group of responses to each question and using within-group relative rewards instead of absolute values, thereby eliminating the critic network. That "relativization" idea runs in exactly the same vein as PPO using advantages instead of absolute returns. See Frontier Progress.

The numbers trap

Never reverse-engineer a conclusion from a paper's numbers. The most common mistake: "DQN stacked 4 frames and it worked, so I'll stack 4 frames too" — that number comes from Atari's pixel input; if your environment has no temporal continuity, stacking 4 frames is pure memory waste. The mechanism is "give the CNN motion information via consecutive frames"; the number is just one instance of that mechanism in one particular setting.

How This Section Relates to the Rest of the Site ​

The papers/ section doesn't stand alone — it forms a closed loop with the whole site: "principle → mechanism → case → paper":

ModuleRelationship to papersHow to use them together
concepts/ (Core Concepts)Papers are the original sources of concepts; concept pages are the "digested versions" of papersStuck on a paper → revisit the concept page; want to trace a concept to its root → jump to the paper
case-studies/ (Case Studies)Papers are the "construction drawings" of the cases; case pages are the "finished renders"Read the AlphaGo case alongside the AlphaGo chapter in Classic Paper Readings
practice/ (Practice)Paper algorithms have to land in codeWhen reproducing a classic paper, pair it with the Progressive Gymnasium Tutorial
guide/ (Guide)The paper pages are the "chain of evidence" behind the history and anatomyEvery decade node in The Evolution of RL points to the paper pages
resources/ (Resources)Provides tooling and terminology supportGlossary /resources/glossary, datasets & tools /resources/datasets-tools

The recommended loop: learn an algorithm on the Core Concepts page → find its original paper on the Paper Map → read the method section in the close readings page → run it once in the practice pages → come back to the paper page for experiments and limitations. An algorithm is truly yours only after you've walked this full loop.

Further Reading ​

  • Reading Paths — three goal-driven reading routes with time budgets; pick your route before you start.
  • Paper Map — seventy years of key RL papers arranged along six branches, for building a global coordinate system.
  • Classic Paper Readings — close readings of six milestone papers: Bellman, Q-learning, DQN, PPO, AlphaGo, InstructGPT.
  • Frontier Progress — frontier directions of the 2020s: world models, offline RL, the RLHF family, RL for reasoning, and more.
  • Reading Discipline & FAQ — methodology: the three-column note-taking method, what to do when the math won't parse, reproduction discipline.
  • Glossary — a quick-reference handbook for unfamiliar terms encountered while reading papers.

References ​

  • Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. Free online edition: https://incompleteideas.net/book/RLbook2020.pdf — the textbook most worth reading cover to cover before diving into papers; its terminology and notation are the standard this site follows.
  • Li, Y. (2017). Deep Reinforcement Learning: An Overview. arXiv:1701.07274 — an entry-level survey, best read to build the overall framework before tackling the classics.
  • Weng, L. (2018). The Policy Gradient Method. https://lilianweng.github.io/posts/2018-04-08-policy-gradient/ — one of the clearest blog explanations of policy gradients from REINFORCE to PPO; a good companion to the close readings.
  • OpenAI Spinning Up in Deep RL: https://spinningup.openai.com/en/latest/ — the official tutorial, with an "abstract-level" explanation of each classic algorithm designed to be read alongside the papers.
  • The References sections of this documentation's Paper Map and Reading Paths pages collect the original links (arXiv, etc.) for every paper mentioned.