Appearance
Reading Discipline and FAQ
In one sentence: this page is the "worksite discipline" for reading RL papers — it answers "how do I take notes, what do I do when the math won't yield, how do I judge whether a paper is worth reading, which papers can I skip, and how do I keep discipline when reproducing experiments." Read it once before tackling any paper (and come back whenever you get stuck); afterwards you'll have a reading workflow of your own instead of grinding by feel.
The previous pages gave you a route (Reading Path), a map (Paper Map), and close-reading samples (Close Readings of Classic Papers). This page fills in the piece most easily missing: method. Most people fail at reading papers not because the papers are hard, but because they have no discipline — no notes, no sense of "how thoroughly should I read this," no verification loop through reproduction. The sections below follow "before reading → while reading → after reading."
1. Reading Discipline: The Three-Column Note Method
Read any paper with notes at hand. Use exactly three columns — no more, or it won't be sustainable:
text
┌────────────────────────────────────────────────────────────────────┐
│ Paper: <title + authors + year + link> │
│ │
│ ① Problem (what are they trying to solve?) │
│ · One-sentence version: │
│ · Why did the previous generation fall short? (what pitfall │
│ did it leave behind) │
│ │
│ ② Method (how do they solve it?) │
│ · Core mechanism in one sentence: │
│ · Key equations / figures / tables (paste or redraw) │
│ · Versus the closest previous paper, what changed? │
│ │
│ ③ Limitations (what did they NOT solve?) │
│ · Assumptions / failures the paper itself admits: │
│ · My own doubts: │
│ · Attackable gaps (fill only if you're doing research): │
└────────────────────────────────────────────────────────────────────┘The golden ratio of the three columns: ① should be short (two or three lines), ② is the main body, and ③ is where you pull ahead of everyone else. Most beginners pour 90% of their time into ② and skip ③ — the result is that they "understood the paper" yet "can't say where it's weak," and one follow-up question in an interview exposes that.
Note-taking tools
Plain text or tables both work (a Markdown table, Notion, Logseq — whatever). The point is a stable entry point for every paper, so that six months later you can recall "what was this paper about" in 30 seconds. If you can fill all three columns for every paper, your reading has reached the L2 mechanism level (see the level table in Reading Path).
2. When the Math Won't Yield: Formula → Intuition → Code
"Give up when the formula makes no sense" is the number-one paper-reading killer. The right order is the three-layer conversion formula → intuition → code — wherever you're stuck, patch that layer:
| Layer | What to do | When stuck |
|---|---|---|
| Formula layer | Understand the symbol definitions and derivation steps | Skip the derivation for now; just pin down "what are the inputs and outputs"; match every symbol in the formula against the paper's notation table |
| Intuition layer | Ask "what intuition is this formula expressing?" | Find a one-sentence explanation on a blog or concept page (this site's concept pages exist exactly for this) |
| Code layer | See how this line is written in the official implementation | This is the ultimate test: if you can read the formula once it becomes fewer than 5 lines of code, you truly understand it |
text
Worked example: when PPO's clipped objective L^CLIP doesn't parse
Stuck at the formula layer → copy down the definition of r_t(θ) (the probability
ratio of new to old policy), and pin down what min() and clip() each do
Intuition layer → "constrain the new-to-old policy ratio within [0.8, 1.2] so a
single update can't move too far" (the PPO section of [Close Readings of
Classic Papers](/papers/core-papers) has a full teardown)
Code layer → open the PPO implementation in SB3/CleanRL, find the two clip lines,
and confirm "clipped_ratio = ratio.clamp(1-eps, 1+eps)" maps one-to-one onto the formulaTwo hard rules: ① if a formula stalls you for more than 30 minutes, skip it and mark it for later; ② never stop at "the formula is copied down" — you must be able to state the intuition in your own words, or you'll forget it within three days. When your math foundation is shaky, the Math Primer — organized as "which RL algorithm does each concept map to" — is a faster way station than a textbook.
3. A Checklist for Judging Paper Quality
Before reading a paper, ask eight questions first; only those answered mostly with "yes" deserve L2/L3 close reading:
| # | Question | Signals of a good paper |
|---|---|---|
| 1 | Is the problem real? | Points to a concrete scenario and its cost (not a vacuous "we propose a novel method") |
| 2 | Is the motivation clear? | States explicitly "previous-generation methods fail in scenario X" and gives examples |
| 3 | Is the method a reusable mechanism? | Distills a transferable idea (e.g., "group-relative advantage"), not just another network-architecture swap |
| 4 | Are the baselines sufficient? | Compares against strong baselines (PPO/SAC/IQL — the de facto standards of each field), and does so fairly |
| 5 | Multiple tasks / multiple seeds? | Reports distributions and variance, not a single best curve |
| 6 | Any ablations? | Breaks down "which mechanism does the work," rather than shipping a black-box bundle |
| 7 | Does it admit limitations? | Has an honest failure-mode discussion (top venues do this best) |
| 8 | Is the code public? | An official implementation exists, lowering your verification cost |
Quick landmine detection
If any of the following shows up, demote the paper straight to L0 (skim the abstract): results reported only on a custom private benchmark; no de facto standard algorithm in the comparison table; missing ablations; conclusions wildly inconsistent with the figures; code "coming soon" for years on end. Your reading time should go to good papers — judging "is this worth reading" is itself part of reading skill.
4. Which Papers You Can Skip
The RL literature numbers in the thousands per year; the ability to skip is worth as much as the ability to read closely. Ranked by skippability:
| Type | Why skippable | Exceptions |
|---|---|---|
| Reskin engineering papers (swap A for B, +1%) | No new mechanism | If the task is your industry's pain point, the experiments section may be worth a look |
| Irreproducible papers (no code + no multi-seed) | High learning cost, zero payoff | None |
| Reviews that state the obvious (someone writes one every year) | The information is already in your head | Top-tier reviews (e.g., annual surveys from authoritative groups) are worth the taxonomy section |
| Single-environment score-chasing papers | Generalization unknown, mechanism unclear | If that environment is your production environment |
| Papers conclusively refuted by later work | The answer is already in the newer work | Just learn about them while reading the refutation |
A more useful principle: it's not "skip the paper," it's "downgrade the reading." No paper is entirely off-limits — most just deserve demotion to L0/L1 (spend 5–15 minutes on the abstract and conclusions). Your time-budget table (Reading Path) tells you which ones deserve promotion to L2+.
5. Reproduction Discipline: Official Implementation First, Your Own Second
Reproduction is the ultimate test of reading, and the easiest place to step on landmines. There is one main line of discipline: always get the official implementation running before writing a line yourself. The order, and the reasons:
text
Step 1: run the official implementation (to get the "right answer")
Purpose: confirm the paper's results hold on your machine (rule out hardware/version differences)
Output: a reproducible run record (seeds, versions, hyperparameters, learning curves)
Step 2: verify generalization on a different environment / dataset
Purpose: distinguish "works in this specific environment" from "the mechanism genuinely works"
Step 3: write it from scratch (without looking at the official code)
Purpose: internalize the mechanism — if you can write down the clip objective,
the replay buffer, and the target network from memory, you've truly read it
(benchmark: the mechanisms in [Close Readings of Classic Papers](/papers/core-papers))
Step 4: ablate
Purpose: remove mechanisms one at a time and watch the impact, answering
"which mechanism is actually doing the work"Common failure modes of these steps and their countermeasures:
| Failure mode | Countermeasure |
|---|---|
| The official implementation can't reproduce the paper's numbers | Check hardware/dependency/hyperparameter differences first; then check whether the paper fixed an unreported seed; only as a last resort suspect "the paper exaggerates" (it does happen) |
| Your rewrite underperforms the official implementation | Diff the two implementations line by line (loss, gradient clipping, entropy regularization, network initialization) — the gap usually hides in the details |
| The reproduction is abandoned halfway | Count "got the official implementation running" as a completion in itself — it's already a solid portfolio project (see Portfolio Projects) |
The minimal reproduction skeleton
Start from the three CartPole versions in the Gymnasium Progressive Tutorial (tabular Q-learning → REINFORCE → PPO). The skeleton is short, readable, and hackable — the best practice ground for reproducing any paper's algorithm.
6. FAQ: A Q&A Wrap-up
Q1: I only have 30 minutes a day — how should I schedule paper reading? Fix a minimal closed loop: 5 minutes to scan 3 new abstracts (decide read or skip) → 20 minutes of close reading on the core mechanism of one L2 paper → 5 minutes writing three-column notes. Don't overestimate the payoff of one deep session, and don't underestimate the compounding of a long-term habit. Five 30-minute sessions a week is 20 papers' worth of mechanism-level accumulation per month.
Q2: What if I can't understand the terminology in an abstract? Check the Glossary first. For words not in the glossary, search the original English term plus a survey of that field. Note: many terms mean different things across papers (e.g., "advantage" is notated differently in RLHF than in classic RL) — when in doubt, trust the usage in the classic papers.
Q3: Should I read blog posts first, or go straight to the paper? Strategy: blog first, paper second. Blogs (e.g., Spinning Up, Lilian Weng) build mechanism-level intuition; the paper verifies the details and the original formulation. But not the other way around — reading only blogs and never papers leaves you missing the "qualifying conditions" (blogs systematically drop the failure discussions). See the "hot topics vs. frontier" analysis in Start Here.
Q4: Do I really have to re-derive the math myself? Depends on your goal. For interviews/engineering: derive the one or two core results (e.g., why PPO's clip is a trust-region surrogate, why Q-learning converges) and trust intuition for the rest. For research: derive the main theorems, and be able to point at "which step of the proof depends on which assumption" — that is where gaps hide. The test: can you answer "if assumption X were dropped, where exactly would the proof break"?
Q5: When should I give up on a paper? Three signals — any one justifies giving up (or downgrading): ① 45 minutes in and you're still spinning at the formula layer, producing no intuition; ② the "limitations" column of your three-column notes stubbornly stays empty (which says the paper — or your reading of it — never reached the mechanism); ③ three weeks of reproduction with no progress and no diagnosable environment issue. Giving up is not failure; it's resource allocation.
Q6: How do I split time between reading papers and running projects? The principle is projects are the backbone, papers are the ribs: let projects generate questions, and let papers answer them. A "reading marathon" of pure input is extremely inefficient; a pure coder who never reads doesn't get far either. See Route 2 in Reading Path: classic papers + one hands-on reproduction is the balance point.
Q7: How do I keep up with the latest papers in a direction? arXiv subscriptions (by keyword: reinforcement learning, RLHF, world models, offline RL), Papers with Code Trending, plus the taxonomy framework in Frontier Advances. Thirty fixed minutes a week for "scanning the new" is plenty — don't let daily updates hold you hostage.
Q8: Does this FAQ apply to all RL papers? Yes. The three-layer formula → intuition → code conversion, the three-column notes, and the official-first reproduction order all hold for the three types of papers — Close Readings of Classic Papers (theory), Paper Map (surveys), and Frontier Advances (the latest). The only difference is "how thoroughly to read each" — let the level table decide.
Further Reading
- Start Here — the section's front door: three routes, and "take away the mechanism, not the numbers."
- Close Readings of Classic Papers — ready-made examples of the three-column method: six papers broken down along "background → method → experiments → limitations" for side-by-side comparison.
- Paper Map — the coordinate system for "which tributary does this paper belong to," a positioning tool to use before reading.
- Gymnasium Progressive Tutorial — the minimal practice ground for reproduction discipline: three CartPole versions from scratch.
References
- Keshav, S. (2007). How to Read a Paper. University of Waterloo. https://web.stanford.edu/class/ee384m/Handouts/HowtoReadPaper.pdf — the classic three-pass methodology (survey / expert-level / reproduction pass); the three-column method in this page was inspired by it.
- OpenAI Spinning Up in Deep RL. https://spinningup.openai.com/en/latest/ — the official tutorial with algorithm-level summaries and reproduction hints; pair it with paper reading.
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. https://incompleteideas.net/book/RLbook2020.pdf — the final authority when a paper's math won't yield: notation and derivations defer to it.
- Weng, L. (2018). The Policy Gradient Method. https://lilianweng.github.io/posts/2018-04-08-policy-gradient/ — a high-quality blog exemplifying the "blog first, paper second" strategy.
- Papers with Code: https://paperswithcode.com/ — find official implementations and check a paper's ranking on benchmarks; before reproducing, come here first to confirm a reference implementation exists.