Skip to content

Learning Paths: Three Routes

On this page These 50+ pages are not a library to read cover to cover but a roadmap — a two-week interview sprint for anyone with interviews in 1–3 months, an ~8-week systematic track for building solid foundations, and a "question → page" desk-reference index for builders already on the job.

Learning Paths: Three Routes ​

One-line pitch: this page is the front door of the site — it turns 50-plus pages from a "library" into a "roadmap": a two-week interview sprint for anyone facing interviews within 1–3 months, an eight-week systematic track for those who want real foundations, and a "question → page" desk-reference index for builders already working in the field. Once you finish this page, you should be able to answer three questions: who am I, where am I headed, and which page should I open right now.

1. The Three Routes at a Glance ​

text
              RL Handbook (50+ pages, six modules)
                               │
       ┌───────────────────────┼───────────────────────┐
       ▼                       ▼                       ▼
  Route 1: Interview      Route 2: Systematic     Route 3: Desk
  Sprint (~2 weeks)       Study (~8 weeks)        Reference (ongoing)
       │                       │                       │
   · career module        · guide → concepts       · question → page
     as the main line      · case studies →         · index
   · high-frequency          practice               · deep dives on
     concepts as sides    · papers → math          · demand
   · 1 Gymnasium demo                                · no required order
       │                       │                       │
       ▼                       ▼                       ▼
   Can answer interview   Can independently        Every debugging
   questions and          build and evaluate       session faster
   explain PPO/RLHF       a real RL project        than the last

The three routes answer three different people's three different questions. The core differences come down to "goal, pace, and how you read":

RouteWho it's forGoalTime budgetHow to read
1. Interview SprintAnyone with interviews in 1–3 monthsAnswer interview questions, talk through projects, survive resume follow-ups~2 weeks (3–4 hours a day)Question-driven reading, efficiency first — memorize the frameworks, then backfill details
2. Systematic StudyBeginners who want to truly understand RL, career changers, studentsIndependently complete an RL project and write up a solid evaluation~8 weeks (10–15 hours a week)Read in order, alternating concepts with hands-on work, with a check to pass every week
3. Desk ReferenceEngineers / researchers in the middle of an RL projectFind an actionable answer fast when a problem hitsNo fixed duration — reused for lifeRead on demand, entering pages from the "problem index" table

One thing not to do

Don't read this site cover to cover like a novel. Reading all 50 pages in sequence is neither realistic nor necessary — most pages stand alone, and which page to read depends entirely on your current goal. That's the whole point of this page: pick a route first, then start reading.

2. Route 1: The Interview Sprint (~2 Weeks) ​

Who it's for: you already have ML / deep learning foundations, and your goal is to land RL-related interviews (algorithm researcher, RL engineer, LLM alignment engineer) within 1–3 months. Your scarce resource is time, not depth of knowledge — so this route runs on one principle: interview questions are the main line, high-frequency concepts are the side dishes, and one hands-on demo is the non-negotiable floor.

1. Main line: the four career-module pages ​

Start with the career module itself — those pages tell you what the market wants and how to benchmark yourself against it:

  1. Read Module Guide & Job Landscape first to map out the responsibilities, skills, and salary bands of 8–10 RL job types, and pick a target position.
  2. Then tear down the JD List: mine real job postings for their high-frequency requirement keywords (policy gradient, PPO, RLHF, simulation, robotics…), and copy them down as your "exam syllabus."
  3. Next, run your resume through Capability Benchmarking: check whether the four elements — environment, algorithm, metrics, reproducibility — are all in place.
  4. As you study, keep cross-checking against the Interview Question Bank, using its answer frameworks to test whether you can turn a concept into an interview answer.

2. Side dishes: four high-frequency concepts you cannot skip ​

Interview questions repeat. The following four concepts show up most often — for each, you must be able to explain it in one sentence plus one example:

ConceptYou must be able to answerRelated pages
MDPWhat's in the five-tuple? Why does the Markov property matter? The intuition for discount factor γ?Markov Decision Process
Value learningMC vs TD? Q-learning vs SARSA? Why does DQN need replay and a target network?Value Learning
Policy gradientThe intuition for the policy gradient theorem? Why is REINFORCE's variance so high? What is PPO's clip doing?Policy Gradient Methods
RLHFThe three-stage pipeline? How is the reward model trained? Why is the KL penalty necessary? What does DPO simplify?RLHF and Alignment with Human Feedback

The interviewer's "explain it to me" test

Talk into a mirror (or a voice recorder) for three minutes on each concept: definition first, then intuition, then an example, then one pitfall. If it flows, you know it. Interviewers' first cut on RL candidates is often "explain the relationship between PPO and TRPO" — fumble that and you're half dead.

3. The hands-on floor: one demo that actually runs ​

An RL candidate with no project on their resume is in danger. This route requires exactly one demo: the Progressive Gymnasium Tutorial (three versions of CartPole: tabular Q-learning → REINFORCE → PPO). It answers three high-frequency interview questions: what does the environment API look like, how do you write a training loop, and how do you read a learning curve.

4. The two-week schedule ​

DaysMain-line taskSide dish / hands-onMilestone
1–2Read Module Guide + JD ListCopy out 15 high-frequency requirement keywordsTarget position chosen; exam syllabus listed
3–4Work through the Interview Question BankRead What Is Reinforcement LearningCan define RL in one sentence
5–6Concept group 1: MDP + value learningGet Gymnasium tutorial version 1 (Q-learning) runningCan write out the Bellman equation from memory
7–8Concept group 2: policy gradient + PPOGet version 2 (REINFORCE) and version 3 (PPO) runningCan explain the clip mechanism
9–10Concept group 3: RLHFRead the LLM Alignment case studyCan draw the three-stage pipeline
11–12Revise your resume with Capability BenchmarkingWrite the demo's evaluation protocol into the READMEResume satisfies the four elements
13–14Full mock: use the Interview Question Bank as an exam paperPatch the gapsEvery high-frequency question gets a 3-minute answer

The sprint's most common trap

Memorizing concepts without building the demo. When the interviewer asks "what RL experiments have you run?" and you have nothing to say, every formula you memorized gets discounted. Even if it's just CartPole, treat it as a real "project" — that's the point the Portfolio Projects page hammers home again and again: one small project that runs, reproduces, and evaluates properly beats ten algorithm names you never ran.

3. Route 2: The Systematic Track (~8 Weeks) ​

Who it's for: learners with no time pressure who want genuine skill that doesn't rely on endorsements. The principle here: guide → core concepts → case studies → practice → papers, one full pass through the five steps — and every week has a "pass bar": clear it before moving on; if not, remediate.

1. Overall rhythm ​

text
Week 1  Orientation & math        ← guide/ module + math-primer
Week 2  MDPs and tabular worlds   ← concepts + hands-on (tabular Q-learning)
Week 3  Value learning            ← concepts + hands-on (MC/TD/DQN ideas)
Week 4  Policy gradients & AC     ← concepts + hands-on (PPO)
Week 5  Deep RL engineering       ← evaluation, frameworks, tuning
Week 6  Classic case studies      ← 3 picks from case-studies
Week 7  Advanced paradigms        ← model-based / offline / multi-agent
Week 8  RLHF and papers           ← alignment + paper map + capstone project

2. Weekly themes and pass bars ​

WeekThemeMust-read pagesMust-doPass bar (before advancing)
1Orientation and math foundationsWhat Is Reinforcement Learning, RL vs Neighboring Paradigms, Math Primer, GlossaryBrowse the full-site Page Map to build a mental overviewCan draw the agent–environment loop; can explain where expectation, conditional expectation, and variance show up in RL
2MDPs and tabular worldsMarkov Decision Process, Multi-Armed Bandits, Exploration vs ExploitationHand-write value iteration or policy iteration on a grid worldCan write out the Bellman expectation equation and optimality equation from memory, and explain why γ<1
3Value learningValue LearningGet tutorial version 1 (tabular Q-learning) running in the Gymnasium TutorialCan discuss MC vs TD bias/variance, and why Q-learning is off-policy
4Policy gradients and Actor-CriticPolicy Gradient Methods, Actor-Critic FamilyGet versions 2 and 3 running; plot a learning curve onceCan derive, not recite, the chain from REINFORCE's high variance → baseline → PPO clip
5Deep RL engineeringEvaluation & Benchmarks, Evaluation in Practice, Frameworks & Tools, Tuning in PracticeBuild an evaluation protocol for your own experiments (multiple seeds)Can design a fair algorithm comparison experiment and write a report template
6Classic case studiesAlphaGo & MCTS, Atari & Video Games, Robotics Sim2RealRe-read the corresponding papersCan explain the "search × learning" complementarity; can explain DQN's two big engineering tricks
7Advanced paradigmsModel-Based RL, Offline RL, Multi-Agent, Reward EngineeringRead one frontier paper (Paper Frontier)Can compare model-based vs model-free, and offline vs online, in terms of when each applies
8RLHF and modern alignmentRLHF, LLM Alignment case study, A Short History of RLDo a close reading of 1–2 papers from Core PapersCan walk through the SFT → reward model → PPO three stages, and explain alignment tax

After week 8

By the end of week 8 you can build things. Two options from here: either go to the Practice module and build a complete project from scratch (environment selection → baseline → training → evaluation → deployment); or move to the Papers module for deep-water reading, picking a line by goal from Paper Reading Paths. At this point Route 3 starts working for you — you already own the map; everything left is on-demand lookup.

4. Route 3: The Desk Reference ​

Who it's for: engineers and researchers in the middle of an RL project, stuck on a specific problem. There's no reading order. The rule is a single line: hit a problem → look it up in the index below → flip to the page → solve it → get on with your day.

Problem you're facingOpen directlyWhat that page gives you
Should this task even use RL?RL vs Neighboring ParadigmsBoundaries across seven paradigms + a selection decision table
Don't know how to cast the problem as RLMarkov Decision ProcessThe five-tuple definition; how to design states/actions
Which simulator to pick / build your own?Datasets & Tools ProfilesA full comparison of environment libraries, benchmarks, datasets
Algorithm won't run / won't learnBuild an RL Project from Scratch + Common PitfallsAn eight-step pipeline + a ten-item failure checklist
The reward keeps getting gamedReward EngineeringA reward hacking case collection + a ten-point design checklist
Want to switch algorithms — PPO or SAC?Actor-Critic Family + Framework ComparisonLineage comparison table + selection advice
How do I tune hyperparameters?Tuning in PracticeCore hyperparameter cheat sheet + tuning order
Results won't reproduceEvaluation in Practice + Evaluation & BenchmarksEvaluation protocols, multi-seed conventions
Only historical data, no more interaction allowedOffline RLThe OOD problem, CQL/IQL, decision trees
Want an LLM to "speak human"RLHF + LLM Alignment case studyThe three-stage pipeline, DPO, alignment tax
Interviews coming up, afraid of questionsInterview Question Bank + JD ListHigh-frequency questions + answer frameworks + topic frequency
Want to see how to read a specific paperCore Papers Close Reading + Paper MapBackground, method, limitations, and how to answer it in interviews

Meta-advice for desk-reference users

Every time you look up a page, spend 30 seconds writing one line in your bookmarks or notes: "problem → conclusion." Three months later you'll own a personal debugging handbook — the only way to turn the reference route into a personal asset. For more systematic note-taking methods, see Paper Reading Discipline & FAQ.

5. The Three Routes Aren't Mutually Exclusive ​

The routes aren't a three-way either/or — they're three ways of opening the same handbook, and you can switch or stack them at any time:

  • Sprint → Study: once the offer lands (or interviews pause), take the concepts you "memorized" during the sprint and reforge them into systematic understanding. The sprint hands you the framework; the systematic track grows that framework into bone — which is why many sprinters find themselves back in the Math Primer by week 2.
  • Study → Reference: after the eight weeks, the systematic track naturally decays into desk-reference mode. You no longer read in order; you just flip to the index when problems hit.
  • Reference → Study: if a reference-only user notices they keep looking up the same kind of question (say, evaluation protocols every time), that's a systemic gap — put that topic back into the weekly plan of Route 2 and run a remedial pass.

All three routes share the same underlying structure: concepts are the bricks, case studies are the house, practice is the construction site, papers are the blueprints, and resources plus career are the tools you bring in. The relationship among the site's six modules fits in one diagram:

text
┌─────────────────────────────────────────────┐
│  guide/   Orientation: start here, pick      │
│           a route                            │
├─────────────────────────────────────────────┤
│  concepts/  Core Concepts (12 pages) ──      │
│           the bricks                         │
│  case-studies/ Case Studies (10 pages) ──    │
│           the house                          │
│  practice/ Practice Guide (8 pages) ──       │
│           the construction site              │
│  papers/  Paper Reading (6 pages) ──         │
│           the blueprints                     │
├─────────────────────────────────────────────┤
│  resources/ Resources: glossary, math,       │
│           checklists, tools                  │
│  career/  Career: landscape, JDs, resumes,   │
│           interviews                         │
└─────────────────────────────────────────────┘

Whichever route you enter by, if you only remember one thing, make it this: hands-on first. RL is a discipline where reading 100 pages is worth less than running 100 steps — and the Gymnasium Tutorial is always your fastest starting point.

Further Reading ​

References ​

  • Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. Full text online: http://incompleteideas.net/book/the-book-2nd.html — the source textbook for every concept on this site.
  • OpenAI (2018). Spinning Up in Deep RL. https://spinningup.openai.com/en/latest/ — a crash course in deep RL aimed at engineers, closely aligned with "Route 1" on this site.
  • Lilian Weng (2017–2025). Lilian Weng's Blog — Reinforcement Learning series. https://lilianweng.github.io/ — high-quality survey-style blog posts, ideal as weekly supplementary reading on the systematic track.
  • Farama Foundation. Gymnasium Documentation. https://gymnasium.farama.org/ — the official documentation for the environment API used in the hands-on demos.
  • David Silver (2015). UCL Course on Reinforcement Learning. https://www.davidsilver.uk/teaching/ — the classic course that maps week-by-week onto weeks 1–4 of "Route 2."