Appearance
Learning Paths: Three Routes
One-line pitch: this page is the front door of the site — it turns 50-plus pages from a "library" into a "roadmap": a two-week interview sprint for anyone facing interviews within 1–3 months, an eight-week systematic track for those who want real foundations, and a "question → page" desk-reference index for builders already working in the field. Once you finish this page, you should be able to answer three questions: who am I, where am I headed, and which page should I open right now.
1. The Three Routes at a Glance
text
RL Handbook (50+ pages, six modules)
│
┌───────────────────────┼───────────────────────┐
▼ ▼ ▼
Route 1: Interview Route 2: Systematic Route 3: Desk
Sprint (~2 weeks) Study (~8 weeks) Reference (ongoing)
│ │ │
· career module · guide → concepts · question → page
as the main line · case studies → · index
· high-frequency practice · deep dives on
concepts as sides · papers → math · demand
· 1 Gymnasium demo · no required order
│ │ │
▼ ▼ ▼
Can answer interview Can independently Every debugging
questions and build and evaluate session faster
explain PPO/RLHF a real RL project than the lastThe three routes answer three different people's three different questions. The core differences come down to "goal, pace, and how you read":
| Route | Who it's for | Goal | Time budget | How to read |
|---|---|---|---|---|
| 1. Interview Sprint | Anyone with interviews in 1–3 months | Answer interview questions, talk through projects, survive resume follow-ups | ~2 weeks (3–4 hours a day) | Question-driven reading, efficiency first — memorize the frameworks, then backfill details |
| 2. Systematic Study | Beginners who want to truly understand RL, career changers, students | Independently complete an RL project and write up a solid evaluation | ~8 weeks (10–15 hours a week) | Read in order, alternating concepts with hands-on work, with a check to pass every week |
| 3. Desk Reference | Engineers / researchers in the middle of an RL project | Find an actionable answer fast when a problem hits | No fixed duration — reused for life | Read on demand, entering pages from the "problem index" table |
One thing not to do
Don't read this site cover to cover like a novel. Reading all 50 pages in sequence is neither realistic nor necessary — most pages stand alone, and which page to read depends entirely on your current goal. That's the whole point of this page: pick a route first, then start reading.
2. Route 1: The Interview Sprint (~2 Weeks)
Who it's for: you already have ML / deep learning foundations, and your goal is to land RL-related interviews (algorithm researcher, RL engineer, LLM alignment engineer) within 1–3 months. Your scarce resource is time, not depth of knowledge — so this route runs on one principle: interview questions are the main line, high-frequency concepts are the side dishes, and one hands-on demo is the non-negotiable floor.
1. Main line: the four career-module pages
Start with the career module itself — those pages tell you what the market wants and how to benchmark yourself against it:
- Read Module Guide & Job Landscape first to map out the responsibilities, skills, and salary bands of 8–10 RL job types, and pick a target position.
- Then tear down the JD List: mine real job postings for their high-frequency requirement keywords (policy gradient, PPO, RLHF, simulation, robotics…), and copy them down as your "exam syllabus."
- Next, run your resume through Capability Benchmarking: check whether the four elements — environment, algorithm, metrics, reproducibility — are all in place.
- As you study, keep cross-checking against the Interview Question Bank, using its answer frameworks to test whether you can turn a concept into an interview answer.
2. Side dishes: four high-frequency concepts you cannot skip
Interview questions repeat. The following four concepts show up most often — for each, you must be able to explain it in one sentence plus one example:
| Concept | You must be able to answer | Related pages |
|---|---|---|
| MDP | What's in the five-tuple? Why does the Markov property matter? The intuition for discount factor γ? | Markov Decision Process |
| Value learning | MC vs TD? Q-learning vs SARSA? Why does DQN need replay and a target network? | Value Learning |
| Policy gradient | The intuition for the policy gradient theorem? Why is REINFORCE's variance so high? What is PPO's clip doing? | Policy Gradient Methods |
| RLHF | The three-stage pipeline? How is the reward model trained? Why is the KL penalty necessary? What does DPO simplify? | RLHF and Alignment with Human Feedback |
The interviewer's "explain it to me" test
Talk into a mirror (or a voice recorder) for three minutes on each concept: definition first, then intuition, then an example, then one pitfall. If it flows, you know it. Interviewers' first cut on RL candidates is often "explain the relationship between PPO and TRPO" — fumble that and you're half dead.
3. The hands-on floor: one demo that actually runs
An RL candidate with no project on their resume is in danger. This route requires exactly one demo: the Progressive Gymnasium Tutorial (three versions of CartPole: tabular Q-learning → REINFORCE → PPO). It answers three high-frequency interview questions: what does the environment API look like, how do you write a training loop, and how do you read a learning curve.
4. The two-week schedule
| Days | Main-line task | Side dish / hands-on | Milestone |
|---|---|---|---|
| 1–2 | Read Module Guide + JD List | Copy out 15 high-frequency requirement keywords | Target position chosen; exam syllabus listed |
| 3–4 | Work through the Interview Question Bank | Read What Is Reinforcement Learning | Can define RL in one sentence |
| 5–6 | Concept group 1: MDP + value learning | Get Gymnasium tutorial version 1 (Q-learning) running | Can write out the Bellman equation from memory |
| 7–8 | Concept group 2: policy gradient + PPO | Get version 2 (REINFORCE) and version 3 (PPO) running | Can explain the clip mechanism |
| 9–10 | Concept group 3: RLHF | Read the LLM Alignment case study | Can draw the three-stage pipeline |
| 11–12 | Revise your resume with Capability Benchmarking | Write the demo's evaluation protocol into the README | Resume satisfies the four elements |
| 13–14 | Full mock: use the Interview Question Bank as an exam paper | Patch the gaps | Every high-frequency question gets a 3-minute answer |
The sprint's most common trap
Memorizing concepts without building the demo. When the interviewer asks "what RL experiments have you run?" and you have nothing to say, every formula you memorized gets discounted. Even if it's just CartPole, treat it as a real "project" — that's the point the Portfolio Projects page hammers home again and again: one small project that runs, reproduces, and evaluates properly beats ten algorithm names you never ran.
3. Route 2: The Systematic Track (~8 Weeks)
Who it's for: learners with no time pressure who want genuine skill that doesn't rely on endorsements. The principle here: guide → core concepts → case studies → practice → papers, one full pass through the five steps — and every week has a "pass bar": clear it before moving on; if not, remediate.
1. Overall rhythm
text
Week 1 Orientation & math ← guide/ module + math-primer
Week 2 MDPs and tabular worlds ← concepts + hands-on (tabular Q-learning)
Week 3 Value learning ← concepts + hands-on (MC/TD/DQN ideas)
Week 4 Policy gradients & AC ← concepts + hands-on (PPO)
Week 5 Deep RL engineering ← evaluation, frameworks, tuning
Week 6 Classic case studies ← 3 picks from case-studies
Week 7 Advanced paradigms ← model-based / offline / multi-agent
Week 8 RLHF and papers ← alignment + paper map + capstone project2. Weekly themes and pass bars
| Week | Theme | Must-read pages | Must-do | Pass bar (before advancing) |
|---|---|---|---|---|
| 1 | Orientation and math foundations | What Is Reinforcement Learning, RL vs Neighboring Paradigms, Math Primer, Glossary | Browse the full-site Page Map to build a mental overview | Can draw the agent–environment loop; can explain where expectation, conditional expectation, and variance show up in RL |
| 2 | MDPs and tabular worlds | Markov Decision Process, Multi-Armed Bandits, Exploration vs Exploitation | Hand-write value iteration or policy iteration on a grid world | Can write out the Bellman expectation equation and optimality equation from memory, and explain why γ<1 |
| 3 | Value learning | Value Learning | Get tutorial version 1 (tabular Q-learning) running in the Gymnasium Tutorial | Can discuss MC vs TD bias/variance, and why Q-learning is off-policy |
| 4 | Policy gradients and Actor-Critic | Policy Gradient Methods, Actor-Critic Family | Get versions 2 and 3 running; plot a learning curve once | Can derive, not recite, the chain from REINFORCE's high variance → baseline → PPO clip |
| 5 | Deep RL engineering | Evaluation & Benchmarks, Evaluation in Practice, Frameworks & Tools, Tuning in Practice | Build an evaluation protocol for your own experiments (multiple seeds) | Can design a fair algorithm comparison experiment and write a report template |
| 6 | Classic case studies | AlphaGo & MCTS, Atari & Video Games, Robotics Sim2Real | Re-read the corresponding papers | Can explain the "search × learning" complementarity; can explain DQN's two big engineering tricks |
| 7 | Advanced paradigms | Model-Based RL, Offline RL, Multi-Agent, Reward Engineering | Read one frontier paper (Paper Frontier) | Can compare model-based vs model-free, and offline vs online, in terms of when each applies |
| 8 | RLHF and modern alignment | RLHF, LLM Alignment case study, A Short History of RL | Do a close reading of 1–2 papers from Core Papers | Can walk through the SFT → reward model → PPO three stages, and explain alignment tax |
After week 8
By the end of week 8 you can build things. Two options from here: either go to the Practice module and build a complete project from scratch (environment selection → baseline → training → evaluation → deployment); or move to the Papers module for deep-water reading, picking a line by goal from Paper Reading Paths. At this point Route 3 starts working for you — you already own the map; everything left is on-demand lookup.
4. Route 3: The Desk Reference
Who it's for: engineers and researchers in the middle of an RL project, stuck on a specific problem. There's no reading order. The rule is a single line: hit a problem → look it up in the index below → flip to the page → solve it → get on with your day.
| Problem you're facing | Open directly | What that page gives you |
|---|---|---|
| Should this task even use RL? | RL vs Neighboring Paradigms | Boundaries across seven paradigms + a selection decision table |
| Don't know how to cast the problem as RL | Markov Decision Process | The five-tuple definition; how to design states/actions |
| Which simulator to pick / build your own? | Datasets & Tools Profiles | A full comparison of environment libraries, benchmarks, datasets |
| Algorithm won't run / won't learn | Build an RL Project from Scratch + Common Pitfalls | An eight-step pipeline + a ten-item failure checklist |
| The reward keeps getting gamed | Reward Engineering | A reward hacking case collection + a ten-point design checklist |
| Want to switch algorithms — PPO or SAC? | Actor-Critic Family + Framework Comparison | Lineage comparison table + selection advice |
| How do I tune hyperparameters? | Tuning in Practice | Core hyperparameter cheat sheet + tuning order |
| Results won't reproduce | Evaluation in Practice + Evaluation & Benchmarks | Evaluation protocols, multi-seed conventions |
| Only historical data, no more interaction allowed | Offline RL | The OOD problem, CQL/IQL, decision trees |
| Want an LLM to "speak human" | RLHF + LLM Alignment case study | The three-stage pipeline, DPO, alignment tax |
| Interviews coming up, afraid of questions | Interview Question Bank + JD List | High-frequency questions + answer frameworks + topic frequency |
| Want to see how to read a specific paper | Core Papers Close Reading + Paper Map | Background, method, limitations, and how to answer it in interviews |
Meta-advice for desk-reference users
Every time you look up a page, spend 30 seconds writing one line in your bookmarks or notes: "problem → conclusion." Three months later you'll own a personal debugging handbook — the only way to turn the reference route into a personal asset. For more systematic note-taking methods, see Paper Reading Discipline & FAQ.
5. The Three Routes Aren't Mutually Exclusive
The routes aren't a three-way either/or — they're three ways of opening the same handbook, and you can switch or stack them at any time:
- Sprint → Study: once the offer lands (or interviews pause), take the concepts you "memorized" during the sprint and reforge them into systematic understanding. The sprint hands you the framework; the systematic track grows that framework into bone — which is why many sprinters find themselves back in the Math Primer by week 2.
- Study → Reference: after the eight weeks, the systematic track naturally decays into desk-reference mode. You no longer read in order; you just flip to the index when problems hit.
- Reference → Study: if a reference-only user notices they keep looking up the same kind of question (say, evaluation protocols every time), that's a systemic gap — put that topic back into the weekly plan of Route 2 and run a remedial pass.
All three routes share the same underlying structure: concepts are the bricks, case studies are the house, practice is the construction site, papers are the blueprints, and resources plus career are the tools you bring in. The relationship among the site's six modules fits in one diagram:
text
┌─────────────────────────────────────────────┐
│ guide/ Orientation: start here, pick │
│ a route │
├─────────────────────────────────────────────┤
│ concepts/ Core Concepts (12 pages) ── │
│ the bricks │
│ case-studies/ Case Studies (10 pages) ── │
│ the house │
│ practice/ Practice Guide (8 pages) ── │
│ the construction site │
│ papers/ Paper Reading (6 pages) ── │
│ the blueprints │
├─────────────────────────────────────────────┤
│ resources/ Resources: glossary, math, │
│ checklists, tools │
│ career/ Career: landscape, JDs, resumes, │
│ interviews │
└─────────────────────────────────────────────┘Whichever route you enter by, if you only remember one thing, make it this: hands-on first. RL is a discipline where reading 100 pages is worth less than running 100 steps — and the Gymnasium Tutorial is always your fastest starting point.
Further Reading
- What Is Reinforcement Learning — read this before entering the systematic track; it lays the foundation for the whole building.
- Markov Decision Process — the universal language of all RL problems, and the mandatory first lesson of every route.
- Anatomy of an RL System — read this when you're ready to switch from "learning concepts" to "building systems," and understand the six-layer architecture.
- A Short History of RL — read this to build a sense of the overall timeline, together with the paper map.
- Core Papers Close Reading — where the sprint route and the systematic track converge in week 8.
- Module Guide & Job Landscape — the sprint route's first stop: look at the market before preparing yourself.
References
- Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. Full text online: http://incompleteideas.net/book/the-book-2nd.html — the source textbook for every concept on this site.
- OpenAI (2018). Spinning Up in Deep RL. https://spinningup.openai.com/en/latest/ — a crash course in deep RL aimed at engineers, closely aligned with "Route 1" on this site.
- Lilian Weng (2017–2025). Lilian Weng's Blog — Reinforcement Learning series. https://lilianweng.github.io/ — high-quality survey-style blog posts, ideal as weekly supplementary reading on the systematic track.
- Farama Foundation. Gymnasium Documentation. https://gymnasium.farama.org/ — the official documentation for the environment API used in the hands-on demos.
- David Silver (2015). UCL Course on Reinforcement Learning. https://www.davidsilver.uk/teaching/ — the classic course that maps week-by-week onto weeks 1–4 of "Route 2."