Appearance
Curated Resource List
In one line: this page is the RL world's "worth paying for (actually, mostly free) list" — textbooks, open courses, in-depth tutorials, blogs, libraries, and channels for tracking papers. Everything on it actually exists and is organized by learning stage, so you can dodge the two classic traps: the "bookmark graveyard" and "whatever a search engine coughs up first, which is usually outdated."
1. Textbooks
| Resource | Author / year | Free? | Positioning | How to read it |
|---|---|---|---|---|
| Reinforcement Learning: An Introduction (2nd ed.) | Sutton & Barto, 2018 | ✅ Free online | The RL bible — the authoritative source for theory + intuition | Read Chapters 1–9 closely; the rest as needed |
| Algorithms for Reinforcement Learning | Csaba Szepesvári, 2010 | ✅ Free online | A tighter, more mathematical treatment (~100 pages) | A math booster once you've seen Sutton & Barto |
| Deep Reinforcement Learning | Aske Plaat, 2022 | ✅ Free online | A panorama of deep RL (open-access via Springer) | Use it to systematize everything from DQN to the frontier |
- Sutton & Barto free version (official): http://incompleteideas.net/book/RLbook2020.pdf
- Szepesvári free version: https://sites.ualberta.ca/~szepesva/papers/RLAlgsInMDPs.pdf
- Plaat free version: https://arxiv.org/abs/2110.14168 (the expanded book version of A Brief Survey of Deep Reinforcement Learning)
TIP
Every chapter of Sutton & Barto ends with exercises and a section on the history and background — the former for self-testing, the latter for understanding "why it was designed this way." Chinese-speaking readers can pair the book with this site's Glossary and Paper Map to keep terminology straight.
2. Open Courses
| Course | Provider | Free? | Strengths | Best for |
|---|---|---|---|---|
| Introduction to Reinforcement Learning (David Silver) | UCL / DeepMind | ✅ | The classic 10-lecture route through tabular methods; the clearest concepts anywhere | Essential for beginners; work through his slides too |
| Deep RL Course (CS285) | UC Berkeley / Sergey Levine | ✅ | A frontier view of deep RL, covering offline RL, meta-RL, and more | Leveling up after the basics |
| DeepMind x UCL RL Lecture Series | DeepMind / UCL (YouTube) | ✅ | The new series from 2022 onward, covering deep RL and the frontier | Your second pass, after Sutton & Barto |
| Deep RL Bootcamp | Berkeley/OpenAI instructor team (2017) | ✅ | Short, punchy lectures with an engineering bent | Getting your hands dirty fast |
- David Silver's course page (videos and slides for every lecture): https://www.davidsilver.uk/teaching/
- CS285 home page (slides and assignments): https://rail.eecs.berkeley.edu/deeprlcourse/
- For the DeepMind x UCL series, search YouTube for "DeepMind x UCL Reinforcement Learning" — the official playlist is on DeepMind's channel.
WARNING
The year of a course matters. David Silver's 2015 course stops at DQN; for deep RL beyond that, go to CS285 or the newer DeepMind series. Don't cite 2015 course material as the state of the art in 2025.
3. In-Depth Tutorials (Hands-On)
| Tutorial | Provider | Strengths | How to get started |
|---|---|---|---|
| Spinning Up in Deep RL | OpenAI | The authoritative intro combining concepts, formulas, and code | Read "Intro to RL" first, then work through the algorithm pages one by one |
| CleanRL | Maintained open-source by vwxyzjn et al. | Minimal, single-file, single-dependency implementations built for teaching | Copy a PPO file and tinker with it |
| Hugging Face Deep RL Course | Hugging Face | Interactive units that run right in Colab/Gradio | Work through the units in order |
| Ray RLlib / SB3 official docs | Ray / Stable-Baselines3 | API and algorithm docs for production-grade tools | Consult them when a project demands it |
- Spinning Up: https://spinningup.openai.com/en/latest/
- CleanRL repo: https://github.com/vwxyzjn/cleanrl ; docs: https://docs.cleanrl.dev/
- Hugging Face Deep RL Course: https://huggingface.co/learn/deep-rl-course/unit0/introduction
How this relates to the site
Our step-by-step tutorial is a homegrown take on the "Spinning Up + CleanRL" formula: it goes from Q-learning all the way to PPO using Gymnasium. Reading it side by side with the CleanRL source code is the fastest way to learn.
4. Blogs and Long-Form Writing
| Blog | Author | What makes it special | Must-reads |
|---|---|---|---|
| Lilian Weng's blog | Lilian Weng | Survey-style long posts covering RLHF, offline RL, and world models | RL from Human Feedback (2023) |
| Distill | The Distill editorial team (on hiatus, but archived) | Interactive visualizations that make mechanisms visible | The classic RL visualization articles |
| The Gradient | An open editorial collective | Opinion pieces and field surveys — good for seeing the debates | The discussions around "The Bitter Lesson" |
| OpenAI / DeepMind official blogs | OpenAI, DeepMind | First-hand explanations when new results drop | The DQN, AlphaGo, and ChatGPT release posts |
| RL Weekly | Community-maintained | A weekly skim of RL papers | Staying current |
- Lilian Weng: https://lilianweng.github.io/
- Distill (archived): https://distill.pub/
- The Gradient: https://thegradient.pub/
- OpenAI blog: https://openai.com/blog ; DeepMind blog: https://deepmind.google/discover/blog/
How to read blogs
Blogs are secondhand sources: they can be outdated or opinionated. The right way to use them: build the map quickly (conclusions first), then go back to the primary papers via the Paper Map. For Lilian Weng's long posts, take notes with the "three-column method" (problem / method / limitations — see Paper Reading Discipline).
5. Tools and Libraries
| Tool | Purpose | Official site | Notes |
|---|---|---|---|
| Gymnasium | The standard environment API (reset/step/render) | https://gymnasium.farama.org/ | The de facto standard; virtually every library is compatible |
| MuJoCo | Physics simulator | https://mujoco.org/ | Free and open source after 2015; the continuous-control benchmark |
| Brax | GPU-accelerated physics simulation | https://github.com/google/brax | Training runs one to two orders of magnitude faster |
| Stable-Baselines3 | Production-grade algorithm library (PyTorch) | https://stable-baselines3.readthedocs.io/ | Works out of the box; good for product validation |
| CleanRL | Minimal implementations for teaching and research | https://github.com/vwxyzjn/cleanrl | The workhorse for reproduction |
| Ray RLlib | Distributed RL | https://docs.ray.io/en/latest/rllib/ | Large-scale parallelism |
| Tianshou | A mid-weight framework with academic flexibility | https://github.com/thu-ml/tianshou | Training loop is highly customizable |
For a fuller comparison, see Choosing Frameworks & Tools; the complete list of environments and datasets lives in Datasets & Tools.
6. Tracking Papers
| Channel | URL | How to use it |
|---|---|---|
| arXiv (cs.LG / cs.AI categories) | https://arxiv.org/ | Set up keyword alerts (e.g. reinforcement learning) |
| Papers with Code | https://paperswithcode.com/ | Papers + official implementations + SOTA leaderboards |
| Google Scholar | https://scholar.google.com/ | Filter for classics by citation count |
| Semantic Scholar | https://www.semanticscholar.org/ | Semantic search + survey recommendations |
| NeurIPS / ICML / ICLR / COLM | The conference sites | The acceptance lists every December / July are a weathervane for the frontier |
TIP
The right way to track papers is not "scroll arXiv every day" but to watch only three to five signals: the Papers with Code RL leaderboards + the Best Paper awards at each NeurIPS/ICML + the citations of two or three senior researchers you trust (Sergey Levine, Pieter Abbeel, Sergey Levine's group, and the like). Read the classics on the Paper Map first, then chase the frontier.
7. Chinese-Language Resources
For Chinese-speaking readers, a few high-quality Chinese-language resources deserve their own list (same bar: everything below actually exists, with format and positioning noted):
| Resource | Format | Positioning | Where to find it |
|---|---|---|---|
| Dive into Reinforcement Learning (Wang Qi, Yang Yiyuan, Jiang Ji) | Open book + companion videos on Bilibili | One of the most complete RL intros in the Chinese-speaking world; classic algorithms implemented in PyTorch | https://hrl.boyuai.com/ |
| Hung-yi Lee's Machine Learning course — the RL lectures | Video course | The most approachable, entertaining RL intro out there; ideal for building intuition the first time | Search YouTube for "Hung-yi Lee" |
| Reinforcement Learning (2nd ed.), Chinese translation | Print book | The Chinese edition of Sutton & Barto, translated by Ye Qiang, Posts & Telecom Press (2019) | Major bookstores |
| Zhou Zhihua's Machine Learning (the "Watermelon Book"), Chapter 16 | Print book | The most concise RL chapter in the Chinese literature, covering tabular methods and policy gradients | Major bookstores |
| Close-reading columns on Zhihu / WeChat official accounts | Long-form articles | Quality varies wildly — use them to look up a concept, never as a textbook | Search for yourself |
Two rules for Chinese-language resources
- Translations disagree on terminology: the same concept shows up under several inconsistent Chinese renderings ("on-policy" alone has at least three common translations). Always work against the original English term — this site's Glossary lists the English originals to help you align.
- Secondhand Chinese content lags: reposted articles on frontier topics (RLHF, offline RL) often arrive six months or more behind the English originals. Chinese resources are fine for getting started and filling gaps, but make frontier judgments from English primary sources.
8. Recommendations by Stage
Everything above, organized into an action list by "where you are right now":
| Stage | Must-do | Optional | What you'll be able to do after |
|---|---|---|---|
| Getting started (months 0–1) | David Silver lectures 1–7 + Spinning Up's Intro to RL + Route 1 of our Learning Paths | The Gymnasium tutorial (ours or the official one) | Explain MDPs, TD, and Q-learning; get CartPole to train |
| Leveling up (months 1–3) | Sutton & Barto Chapters 1–9 + CS285 lectures 1–8 + reimplement PPO from CleanRL | Lilian Weng's RL from Human Feedback post, RLHF-related material | Read the formulas in mainstream papers; reproduce DQN/PPO/SAC |
| Research / engineering frontier (month 3+) | The back half of CS285 + close-read 10–15 classics from the Paper Map + reproduce results from Papers with Code | Follow RL Weekly + conference papers | Form independent judgments; read new papers critically |
Pitfall warning
Plenty of RL tutorials from 2016–2018 still rank in search results, but their content is stale — still presenting "DQN is SOTA" or treating Atari as the main battleground. The test is simple: does the tutorial's idea of SOTA consist of post-2020 algorithms (PPO/SAC/Dreamer/IQL and friends)? If not, treat it as history.
9. Community and Continuous Learning
Resources aren't enough — you also need to live in the field. RL moves fast, and the community is the best source of signal:
| Channel | URL | What it's for |
|---|---|---|
| Reddit r/reinforcementlearning | https://www.reddit.com/r/reinforcementlearning/ | Industry/academia discussion, resource recommendations, help |
| Awesome RL list (aikorea/awesome-rl) | https://github.com/aikorea/awesome-rl | A big, categorized index of RL resources |
| NeurIPS / ICML / ICLR | https://neurips.cc/ , https://icml.cc/ | Two concentrated waves of frontier papers every year |
| arXiv cs.LG | https://arxiv.org/list/cs.LG/recent | Daily new papers (pair with keyword subscriptions) |
| Academic Twitter / researchers' homepages | Levine, Schulman, Lilian Weng, etc. | The fastest frontier commentary and work-in-progress updates |
The 30-second checklist: is a resource worth your time?
Before bookmarking any tutorial, blog, or repo, spend 30 seconds on these:
- Date: updated within the last year or two? RL turns over a generation of algorithms annually.
- Baseline: does it benchmark against algorithms that are still mainstream (the PPO/SAC/Dreamer/IQL generation)?
- Reproducibility: is there code? Does it actually run (check whether Issues go unanswered for ages)?
- Provenance: can you verify the author (affiliation, GitHub history)? Or is it a content farm?
- Is it a rehash?: does it mirror the structure of Spinning Up / CleanRL with nothing added? Then it's secondhand — go to the source and save yourself the time.
How this site helps
Our Learning Paths weave these external resources into three routes, and the Paper Map upgrades you from "consuming resources" to "reading papers." External resources are the raw material, and these pages are the map — they work best together.
10. Turning Resources into Ability: an Action Template
Bookmarking ≠ learning. The point of a resource list is that it gets used, so here is a four-week rhythm template that has been tested in practice (it drops straight onto the systematic track in Learning Paths):
| Week | Theme | Resource combo | Acceptance test |
|---|---|---|---|
| Week 1 | MDPs and tabular methods | David Silver lectures 1–4 + Sutton & Barto Chapters 3–4 + our MDP page | Write the Bellman equation from memory; write down the Q-learning update |
| Week 2 | Value-based learning | Sutton & Barto Chapters 5–6 + Spinning Up's Intro to RL | Explain the bias–variance difference between MC and TD; get Q-learning running on CartPole |
| Week 3 | Policy gradients | CS285 lectures 4–5 + CleanRL's PPO implementation | Explain every term of PPO's clipped objective; train PPO and plot a learning curve |
| Week 4 | End-to-end reproduction | Pick one classic paper from the Paper Map and reproduce it | Ship a README covering the four elements: environment – algorithm – evaluation – reproduction |
Three rules of engagement:
- Output immediately after input: go through each resource exactly once, then immediately write notes / run code / explain it to someone. Hoarding courses you'll "watch later" is banned.
- Judge by results: finishing a course means "produced your own results," not "watched all the videos."
- Cap concurrent resources: at most two courses and one textbook at a time. Running many resources in parallel is usually procrastination in disguise, not diligence.
Frequently asked questions (FAQ)
- Q: I've bookmarked a pile of courses but can't bring myself to start? A: Start with the "smallest actionable step" — watch one lecture today, or just get CartPole running. Don't draft a "month-long plan" as a way to postpone.
- Q: There are too many resources — how do I know which to prioritize? A: Find your stage in the table above and start with the "Must-do" column; the pitfall checklist in Section 7 will filter out the stale content.
- Q: English resources are a struggle — what do I do? A: Start with the Chinese-language resources in Section 7 to build a base, then go back to the English originals; use this site's Glossary for terminology.
How this relates to the site
External resources feed you knowledge; these pages build structure. After each external resource, come back to the matching concept page and check it against the site — "where do they disagree?" That kind of active processing is the most effective way to remember anything.
11. Maintaining Your Own Resource List
Bookmarks rot — links die, courses go offline, versions drift. Three rules turn "someone else's list" into "your list":
- Clean house every quarter: delete anything you haven't opened in three months. If a resource wasn't worth opening for three months, it probably never will be.
- Tag every item with a status: to watch / watching / absorbed / outdated. The status changes themselves are a sense of progress.
- Record where you learned it: whenever a new concept finally clicks, note the first place you understood it (a blog post, a lecture, a page on this site). When it's time to review, go back to the source you understood best.
A ready-to-use list template (plain Markdown is fine — no tool required):
markdown
## Currently learning
- [Course name] (3/10 lectures done) → from the [Curated Resource List](/resources/awesome), Section X
- [Book title], Chapter 5
## Absorbed (cite anytime)
- Concept X: cross-checked against the [Glossary](/resources/glossary); see [Page name](/concepts/value-based)
## Outdated / retired
- [Old tutorial] (2021, treats DQN as SOTA)Resources are a means, not an end: what you still remember a month later, can explain to someone else, and can produce results with — that is what you truly own. The rest of this site — Learning Paths, the Paper Map, and Frameworks & Tools — is here to stay, helping you turn resources into ability.
Further Reading
- Learning Paths: Three Routes — how this site organizes the resources on this page into "what to learn first, what to learn next."
- Glossary — look up unfamiliar terms from these resources as you go.
- Datasets & Tools — the full version of this page's tool list (environment libraries, benchmarks, offline datasets included).
- Choosing Frameworks & Tools — detailed comparisons and selection advice for six frameworks.
- Paper Map — the place to start when you're ready to turn "reading resources" into a system.
References
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd ed. Free online: http://incompleteideas.net/book/RLbook2020.pdf
- David Silver's UCL course page: https://www.davidsilver.uk/teaching/
- UC Berkeley CS285 Deep RL: https://rail.eecs.berkeley.edu/deeprlcourse/
- OpenAI Spinning Up: https://spinningup.openai.com/en/latest/
- CleanRL: https://github.com/vwxyzjn/cleanrl
- Hugging Face Deep RL Course: https://huggingface.co/learn/deep-rl-course/unit0/introduction
- Lilian Weng's blog: https://lilianweng.github.io/
- Gymnasium: https://gymnasium.farama.org/
- Stable-Baselines3: https://stable-baselines3.readthedocs.io/
- Brax: https://github.com/google/brax
- arXiv: https://arxiv.org/ ; Papers with Code: https://paperswithcode.com/