Skip to content

Curated Resource List

On this page Textbooks (Sutton & Barto), courses (DeepMind/Berkeley/UCL), tutorials (Spinning Up, CleanRL), blogs (Lilian Weng), podcasts and communities — a genuinely real resource list organized by learning stage.

Curated Resource List ​

In one line: this page is the RL world's "worth paying for (actually, mostly free) list" — textbooks, open courses, in-depth tutorials, blogs, libraries, and channels for tracking papers. Everything on it actually exists and is organized by learning stage, so you can dodge the two classic traps: the "bookmark graveyard" and "whatever a search engine coughs up first, which is usually outdated."

1. Textbooks ​

ResourceAuthor / yearFree?PositioningHow to read it
Reinforcement Learning: An Introduction (2nd ed.)Sutton & Barto, 2018✅ Free onlineThe RL bible — the authoritative source for theory + intuitionRead Chapters 1–9 closely; the rest as needed
Algorithms for Reinforcement LearningCsaba Szepesvári, 2010✅ Free onlineA tighter, more mathematical treatment (~100 pages)A math booster once you've seen Sutton & Barto
Deep Reinforcement LearningAske Plaat, 2022✅ Free onlineA panorama of deep RL (open-access via Springer)Use it to systematize everything from DQN to the frontier

TIP

Every chapter of Sutton & Barto ends with exercises and a section on the history and background — the former for self-testing, the latter for understanding "why it was designed this way." Chinese-speaking readers can pair the book with this site's Glossary and Paper Map to keep terminology straight.

2. Open Courses ​

CourseProviderFree?StrengthsBest for
Introduction to Reinforcement Learning (David Silver)UCL / DeepMind✅The classic 10-lecture route through tabular methods; the clearest concepts anywhereEssential for beginners; work through his slides too
Deep RL Course (CS285)UC Berkeley / Sergey Levine✅A frontier view of deep RL, covering offline RL, meta-RL, and moreLeveling up after the basics
DeepMind x UCL RL Lecture SeriesDeepMind / UCL (YouTube)✅The new series from 2022 onward, covering deep RL and the frontierYour second pass, after Sutton & Barto
Deep RL BootcampBerkeley/OpenAI instructor team (2017)✅Short, punchy lectures with an engineering bentGetting your hands dirty fast

WARNING

The year of a course matters. David Silver's 2015 course stops at DQN; for deep RL beyond that, go to CS285 or the newer DeepMind series. Don't cite 2015 course material as the state of the art in 2025.

3. In-Depth Tutorials (Hands-On) ​

TutorialProviderStrengthsHow to get started
Spinning Up in Deep RLOpenAIThe authoritative intro combining concepts, formulas, and codeRead "Intro to RL" first, then work through the algorithm pages one by one
CleanRLMaintained open-source by vwxyzjn et al.Minimal, single-file, single-dependency implementations built for teachingCopy a PPO file and tinker with it
Hugging Face Deep RL CourseHugging FaceInteractive units that run right in Colab/GradioWork through the units in order
Ray RLlib / SB3 official docsRay / Stable-Baselines3API and algorithm docs for production-grade toolsConsult them when a project demands it

How this relates to the site

Our step-by-step tutorial is a homegrown take on the "Spinning Up + CleanRL" formula: it goes from Q-learning all the way to PPO using Gymnasium. Reading it side by side with the CleanRL source code is the fastest way to learn.

4. Blogs and Long-Form Writing ​

BlogAuthorWhat makes it specialMust-reads
Lilian Weng's blogLilian WengSurvey-style long posts covering RLHF, offline RL, and world modelsRL from Human Feedback (2023)
DistillThe Distill editorial team (on hiatus, but archived)Interactive visualizations that make mechanisms visibleThe classic RL visualization articles
The GradientAn open editorial collectiveOpinion pieces and field surveys — good for seeing the debatesThe discussions around "The Bitter Lesson"
OpenAI / DeepMind official blogsOpenAI, DeepMindFirst-hand explanations when new results dropThe DQN, AlphaGo, and ChatGPT release posts
RL WeeklyCommunity-maintainedA weekly skim of RL papersStaying current

How to read blogs

Blogs are secondhand sources: they can be outdated or opinionated. The right way to use them: build the map quickly (conclusions first), then go back to the primary papers via the Paper Map. For Lilian Weng's long posts, take notes with the "three-column method" (problem / method / limitations — see Paper Reading Discipline).

5. Tools and Libraries ​

ToolPurposeOfficial siteNotes
GymnasiumThe standard environment API (reset/step/render)https://gymnasium.farama.org/The de facto standard; virtually every library is compatible
MuJoCoPhysics simulatorhttps://mujoco.org/Free and open source after 2015; the continuous-control benchmark
BraxGPU-accelerated physics simulationhttps://github.com/google/braxTraining runs one to two orders of magnitude faster
Stable-Baselines3Production-grade algorithm library (PyTorch)https://stable-baselines3.readthedocs.io/Works out of the box; good for product validation
CleanRLMinimal implementations for teaching and researchhttps://github.com/vwxyzjn/cleanrlThe workhorse for reproduction
Ray RLlibDistributed RLhttps://docs.ray.io/en/latest/rllib/Large-scale parallelism
TianshouA mid-weight framework with academic flexibilityhttps://github.com/thu-ml/tianshouTraining loop is highly customizable

For a fuller comparison, see Choosing Frameworks & Tools; the complete list of environments and datasets lives in Datasets & Tools.

6. Tracking Papers ​

ChannelURLHow to use it
arXiv (cs.LG / cs.AI categories)https://arxiv.org/Set up keyword alerts (e.g. reinforcement learning)
Papers with Codehttps://paperswithcode.com/Papers + official implementations + SOTA leaderboards
Google Scholarhttps://scholar.google.com/Filter for classics by citation count
Semantic Scholarhttps://www.semanticscholar.org/Semantic search + survey recommendations
NeurIPS / ICML / ICLR / COLMThe conference sitesThe acceptance lists every December / July are a weathervane for the frontier

TIP

The right way to track papers is not "scroll arXiv every day" but to watch only three to five signals: the Papers with Code RL leaderboards + the Best Paper awards at each NeurIPS/ICML + the citations of two or three senior researchers you trust (Sergey Levine, Pieter Abbeel, Sergey Levine's group, and the like). Read the classics on the Paper Map first, then chase the frontier.

7. Chinese-Language Resources ​

For Chinese-speaking readers, a few high-quality Chinese-language resources deserve their own list (same bar: everything below actually exists, with format and positioning noted):

ResourceFormatPositioningWhere to find it
Dive into Reinforcement Learning (Wang Qi, Yang Yiyuan, Jiang Ji)Open book + companion videos on BilibiliOne of the most complete RL intros in the Chinese-speaking world; classic algorithms implemented in PyTorchhttps://hrl.boyuai.com/
Hung-yi Lee's Machine Learning course — the RL lecturesVideo courseThe most approachable, entertaining RL intro out there; ideal for building intuition the first timeSearch YouTube for "Hung-yi Lee"
Reinforcement Learning (2nd ed.), Chinese translationPrint bookThe Chinese edition of Sutton & Barto, translated by Ye Qiang, Posts & Telecom Press (2019)Major bookstores
Zhou Zhihua's Machine Learning (the "Watermelon Book"), Chapter 16Print bookThe most concise RL chapter in the Chinese literature, covering tabular methods and policy gradientsMajor bookstores
Close-reading columns on Zhihu / WeChat official accountsLong-form articlesQuality varies wildly — use them to look up a concept, never as a textbookSearch for yourself

Two rules for Chinese-language resources

  1. Translations disagree on terminology: the same concept shows up under several inconsistent Chinese renderings ("on-policy" alone has at least three common translations). Always work against the original English term — this site's Glossary lists the English originals to help you align.
  2. Secondhand Chinese content lags: reposted articles on frontier topics (RLHF, offline RL) often arrive six months or more behind the English originals. Chinese resources are fine for getting started and filling gaps, but make frontier judgments from English primary sources.

8. Recommendations by Stage ​

Everything above, organized into an action list by "where you are right now":

StageMust-doOptionalWhat you'll be able to do after
Getting started (months 0–1)David Silver lectures 1–7 + Spinning Up's Intro to RL + Route 1 of our Learning PathsThe Gymnasium tutorial (ours or the official one)Explain MDPs, TD, and Q-learning; get CartPole to train
Leveling up (months 1–3)Sutton & Barto Chapters 1–9 + CS285 lectures 1–8 + reimplement PPO from CleanRLLilian Weng's RL from Human Feedback post, RLHF-related materialRead the formulas in mainstream papers; reproduce DQN/PPO/SAC
Research / engineering frontier (month 3+)The back half of CS285 + close-read 10–15 classics from the Paper Map + reproduce results from Papers with CodeFollow RL Weekly + conference papersForm independent judgments; read new papers critically

Pitfall warning

Plenty of RL tutorials from 2016–2018 still rank in search results, but their content is stale — still presenting "DQN is SOTA" or treating Atari as the main battleground. The test is simple: does the tutorial's idea of SOTA consist of post-2020 algorithms (PPO/SAC/Dreamer/IQL and friends)? If not, treat it as history.

9. Community and Continuous Learning ​

Resources aren't enough — you also need to live in the field. RL moves fast, and the community is the best source of signal:

ChannelURLWhat it's for
Reddit r/reinforcementlearninghttps://www.reddit.com/r/reinforcementlearning/Industry/academia discussion, resource recommendations, help
Awesome RL list (aikorea/awesome-rl)https://github.com/aikorea/awesome-rlA big, categorized index of RL resources
NeurIPS / ICML / ICLRhttps://neurips.cc/ , https://icml.cc/Two concentrated waves of frontier papers every year
arXiv cs.LGhttps://arxiv.org/list/cs.LG/recentDaily new papers (pair with keyword subscriptions)
Academic Twitter / researchers' homepagesLevine, Schulman, Lilian Weng, etc.The fastest frontier commentary and work-in-progress updates

The 30-second checklist: is a resource worth your time? ​

Before bookmarking any tutorial, blog, or repo, spend 30 seconds on these:

  1. Date: updated within the last year or two? RL turns over a generation of algorithms annually.
  2. Baseline: does it benchmark against algorithms that are still mainstream (the PPO/SAC/Dreamer/IQL generation)?
  3. Reproducibility: is there code? Does it actually run (check whether Issues go unanswered for ages)?
  4. Provenance: can you verify the author (affiliation, GitHub history)? Or is it a content farm?
  5. Is it a rehash?: does it mirror the structure of Spinning Up / CleanRL with nothing added? Then it's secondhand — go to the source and save yourself the time.

How this site helps

Our Learning Paths weave these external resources into three routes, and the Paper Map upgrades you from "consuming resources" to "reading papers." External resources are the raw material, and these pages are the map — they work best together.

10. Turning Resources into Ability: an Action Template ​

Bookmarking ≠ learning. The point of a resource list is that it gets used, so here is a four-week rhythm template that has been tested in practice (it drops straight onto the systematic track in Learning Paths):

WeekThemeResource comboAcceptance test
Week 1MDPs and tabular methodsDavid Silver lectures 1–4 + Sutton & Barto Chapters 3–4 + our MDP pageWrite the Bellman equation from memory; write down the Q-learning update
Week 2Value-based learningSutton & Barto Chapters 5–6 + Spinning Up's Intro to RLExplain the bias–variance difference between MC and TD; get Q-learning running on CartPole
Week 3Policy gradientsCS285 lectures 4–5 + CleanRL's PPO implementationExplain every term of PPO's clipped objective; train PPO and plot a learning curve
Week 4End-to-end reproductionPick one classic paper from the Paper Map and reproduce itShip a README covering the four elements: environment – algorithm – evaluation – reproduction

Three rules of engagement:

  1. Output immediately after input: go through each resource exactly once, then immediately write notes / run code / explain it to someone. Hoarding courses you'll "watch later" is banned.
  2. Judge by results: finishing a course means "produced your own results," not "watched all the videos."
  3. Cap concurrent resources: at most two courses and one textbook at a time. Running many resources in parallel is usually procrastination in disguise, not diligence.

Frequently asked questions (FAQ) ​

  • Q: I've bookmarked a pile of courses but can't bring myself to start? A: Start with the "smallest actionable step" — watch one lecture today, or just get CartPole running. Don't draft a "month-long plan" as a way to postpone.
  • Q: There are too many resources — how do I know which to prioritize? A: Find your stage in the table above and start with the "Must-do" column; the pitfall checklist in Section 7 will filter out the stale content.
  • Q: English resources are a struggle — what do I do? A: Start with the Chinese-language resources in Section 7 to build a base, then go back to the English originals; use this site's Glossary for terminology.

How this relates to the site

External resources feed you knowledge; these pages build structure. After each external resource, come back to the matching concept page and check it against the site — "where do they disagree?" That kind of active processing is the most effective way to remember anything.

11. Maintaining Your Own Resource List ​

Bookmarks rot — links die, courses go offline, versions drift. Three rules turn "someone else's list" into "your list":

  1. Clean house every quarter: delete anything you haven't opened in three months. If a resource wasn't worth opening for three months, it probably never will be.
  2. Tag every item with a status: to watch / watching / absorbed / outdated. The status changes themselves are a sense of progress.
  3. Record where you learned it: whenever a new concept finally clicks, note the first place you understood it (a blog post, a lecture, a page on this site). When it's time to review, go back to the source you understood best.

A ready-to-use list template (plain Markdown is fine — no tool required):

markdown
## Currently learning
- [Course name] (3/10 lectures done) → from the [Curated Resource List](/resources/awesome), Section X
- [Book title], Chapter 5

## Absorbed (cite anytime)
- Concept X: cross-checked against the [Glossary](/resources/glossary); see [Page name](/concepts/value-based)

## Outdated / retired
- [Old tutorial] (2021, treats DQN as SOTA)

Resources are a means, not an end: what you still remember a month later, can explain to someone else, and can produce results with — that is what you truly own. The rest of this site — Learning Paths, the Paper Map, and Frameworks & Tools — is here to stay, helping you turn resources into ability.

Further Reading ​

  • Learning Paths: Three Routes — how this site organizes the resources on this page into "what to learn first, what to learn next."
  • Glossary — look up unfamiliar terms from these resources as you go.
  • Datasets & Tools — the full version of this page's tool list (environment libraries, benchmarks, offline datasets included).
  • Choosing Frameworks & Tools — detailed comparisons and selection advice for six frameworks.
  • Paper Map — the place to start when you're ready to turn "reading resources" into a system.

References ​