Skip to content

REINFORCEMENT LEARNING HANDBOOK

Reinforcement Learning Handbook

A systematic path from trial-and-error to decision-making — MDPs · Value Learning · Policy Gradients · Actor-Critic · Offline RL · Multi-Agent · RLHF · Career Paths

If supervised learning learns to judge and unsupervised learning learns structure, reinforcement learning learns sequences of decisions — an agent learns long-run optimal policies through trial-and-error with its environment, from sparse reward signals alone. Learn more →

Where to Start

50+ pages aren't a library to read cover to cover — they're routes you combine as needed

✨ Editors' Picks

If you only read ten pages, read these

All Content

Browse by module, or search directly