Appearance
What Is Reinforcement Learning
One-line pitch: this page pins down exactly what the words "reinforcement learning" mean — what problem it solves, what interaction framework it uses, how it differs from the supervised and unsupervised learning you already know, and what a minimal example looks like. By the end, you'll be able to tell anyone: RL is the machine learning paradigm where an agent learns what to do by trial and error in an environment, guided by rewards that arrive after the fact.
1. A One-Sentence Definition and the Agent–Environment Loop
Reinforcement Learning (RL) studies how an agent learns a policy through repeated interaction with an environment, so as to maximize the cumulative reward accumulated over the long run.
The catch comes down to three things RL does not have:
- No labels: nobody tells you "the correct action right now is a*";
- No immediate feedback: the consequences of an action — good or bad — may only show up much later;
- No independent and identically distributed data: the state you see next depends on what you just did — the data is generated, and contaminated, by your own behavior.
That is the deepest difference from supervised learning: supervised learning learns "what is this"; RL learns "what should I do".
The interaction framework fits in one diagram:
text
┌────────────────────────────┐
│ Environment │
│ (rules, simulator, │
│ opponents, users) │
└────────────────────────────┘
▲ │
state s_t │ │ action a_t
(+ reward r_t)│ ▼
┌────────────────────────────┐
│ Agent │
│ (policy π, value function,│
│ memory) │
└────────────────────────────┘What happens at each time step t is, mathematically, one sample from a Markov decision process (MDP) — see Markov Decision Process for the full definition. For now, the plain-spoken version:
text
Loop:
1. The environment is in state s_t and hands s_t (plus reward r_t) to the agent
2. The agent picks action a_t according to policy π(a|s)
3. The environment produces the next state s_{t+1} via transition function P(s_{t+1} | s_t, a_t)
4. The environment produces the immediate reward r_{t+1} via reward function R(s_t, a_t)
5. Go back to 1, until a terminal state (or the budget runs out)If you only read one thing, read this
The loop above is the skeleton of the entire building. Every algorithm that follows — Q-learning, DQN, PPO, SAC, RLHF — is a different answer to the same question: "how do I update the agent's internal parameters inside this loop?" Burn this diagram into your brain; when you later read Anatomy of an RL System, you'll watch it scale up into a six-layer engineering system.
2. The Reward Hypothesis and the Four Elements
1. The Reward Hypothesis
RL rests on an axiom known as the reward hypothesis, stated explicitly by Sutton & Barto:
Any goal can be formulated as the maximization of the expected value of a scalar reward signal (cumulative reward).
"Live a long, healthy life" becomes "+1 for every extra day alive." "Win a game of Go" becomes "+1 for a win, −1 for a loss." "Recommend content users like" becomes "+0.1 per click, +0.5 per purchase." Whether this translation is any good decides whether your RL project lives or dies — which is the entire reason the Reward Engineering page exists.
The reward hypothesis is a hypothesis
Translating a goal into a reward is almost always lossy: get the translation wrong, and the agent will game the reward instead of doing what you want (reward hacking). The classic example: a cleaning robot discovers it earns more cleaning reward by dirtying the rug and then cleaning it. This isn't a bug — it's a reward design problem. See the case collection in Reward Engineering.
2. The Four Elements: State, Action, Policy, Return
| Element | English | Definition | An intuition (CartPole) |
|---|---|---|---|
| State | State s | Everything the environment exposes to the agent at a given moment | Cart position and velocity, pole angle and angular velocity |
| Action | Action a | One choice the agent can make | Push left, push right |
| Policy | Policy π(a|s) | A mapping from states to actions (or distributions over actions) | Pole tilts far left → push left |
| Return | Return G_t | The (discounted) sum of rewards from time t onward | The total of "not falling" rewards in the future |
The definition of return (with discount factor γ):
text
G_t = r_{t+1} + γ·r_{t+2} + γ²·r_{t+3} + ... = Σ_{k=0}^{∞} γ^k · r_{t+k+1}The discount factor γ ∈ [0, 1) does three jobs:
- Mathematically, it makes the infinite sum converge;
- It encodes "a dollar today is worth more than a dollar tomorrow" — decision-makers naturally prefer earlier rewards;
- It controls the agent's "foresight": the closer γ is to 1, the more long-sighted the agent; the smaller γ, the more myopic.
Why γ ≠ 1 is the sensible choice
If γ = 1 and the task never ends, the return can diverge; even when it doesn't, "one dollar today equals one dollar ten years from now" defies common sense. In most engineering settings, γ sits between 0.9 and 0.99. γ is the first hyperparameter you'll touch in Tuning in Practice.
3. A Three-Way Comparison with Supervised and Unsupervised Learning
RL is often called the "third paradigm" of machine learning. Here's a side-by-side look along three dimensions:
| Dimension | Supervised learning | Unsupervised learning | Reinforcement learning |
|---|---|---|---|
| Form of data | Input–label pairs (x, y) | Inputs x only | State–action–reward trajectories (s, a, r) sequences |
| Feedback signal | Immediate, explicit label for every sample | No labels; relies on intrinsic structure of data | After-the-fact, sparse, delayed scalar rewards |
| Learning goal | Fit an input→output mapping | Discover distributions/structure (clustering, compression) | Maximize long-term cumulative return through interaction |
| Who produces the data | Given externally (a fixed dataset) | Given externally | The agent's own behavior generates the follow-up data |
| Can errors be corrected immediately? | Yes — the loss hands you gradients directly | No clear right or wrong | No — an action may have to "take the blame" for outcomes long afterward |
| Typical tasks | Classification, regression, detection | Clustering, dimensionality reduction, generation | Games, control, dialogue, ranking, alignment |
| Success metric | Test-set accuracy | Cluster quality / reconstruction error | Trajectory return, task success rate, sample efficiency |
The key difference boils down to one sentence: supervised learning's data is stationary; RL's data is dynamic — you are changing the very distribution you will learn from. This distribution shift (non-stationarity) is the root of all RL difficulty, and it explains why RL needs exploration and why offline data can serve as a safety net.
Another angle on the same idea
Think of supervised learning as a course where the teacher grades every homework assignment, and RL as a course where you only get a final grade and nothing is ever graded in between — and where each assignment you submit changes what the next class covers. To sharpen the boundaries further (RL vs optimal control, RL vs behavior cloning), head straight to RL vs Neighboring Paradigms.
4. Why "Sequences of Decisions" Are Hard: Credit Assignment and Delayed Rewards
A single decision is easy: evaluate the immediate reward of each option and pick the largest. RL, however, deals with sequential decision-making — an action you take today may pay off a thousand steps later. This raises two fundamental problems.
1. The Credit Assignment Problem
When the final reward finally arrives, which of your past actions should get the credit?
text
Scene: after 10 moves of chess, move 3 plants a seed,
and move 10 wins the game.
Question: who gets the +1? Move 3? Or move 10?Credit only the last move, and the agent never learns to "set things up"; spread it evenly across every move, and move 3's contribution gets diluted. TD learning, eligibility traces, GAE, value networks — half of all RL algorithms are answers to this very question.
2. Delayed and Sparse Rewards
Many real-world tasks offer extremely sparse rewards: a robot may take ten thousand steps before hitting a single reward (grasping an object, reaching a goal). With rewards that rare, random exploration almost never stumbles onto one, and learning stalls. Remedies include reward shaping, curriculum learning, and intrinsic rewards (curiosity) — see Exploration vs Exploitation and Reward Engineering for details.
A common beginner's mistake
Assuming that "designing a dense reward will fix it." A dense reward does ease sparsity, but it brings reward hacking and the risk of "locally optimal reward-farmers" (a robot that learns to spin in place farming reward, for instance). Reward density is not the higher the better — it's a trade-off between density and resistance to being gamed.
3. Why Look-Up Tables Fail in the Real World
Credit assignment + sparse rewards + high-dimensional states: three mountains that crush any "table lookup" method. Real state spaces are continuous and vast (images, sensor readings, text) — you cannot store a value per state. The fix is function approximation: use a neural network to learn the mapping from "state → value/action distribution." That is exactly what the two main lines — Value Learning and Policy Gradient Methods — each set out to do.
5. A Minimal RL Example: Q-Learning in a Grid World
Enough theory — let's dissect a minimal example end to end: a 3×3 grid world.
text
┌────┬────┬────┐
│ S │ · │ · │ S = start
├────┼────┼────┤ G = goal (reward +10, terminate)
│ · │ ✗ │ · │ ✗ = trap (reward -5, terminate)
├────┼────┼────┤ other cells: -0.1 per step (hurry up!)
│ · │ · │ G │ actions: up/down/left/right
└────┴────┴────┘The agent doesn't know the map; it can only learn by "taking a step and seeing the reward." We solve it with Q-learning — a Q table storing "the long-term value of doing each action in each cell":
python
import random
# states: 9 cells numbered 0~8; actions: 0=up 1=down 2=left 3=right
gamma = 0.9 # discount factor: closer to 1 = more foresight
alpha = 0.1 # learning rate: how fast new info replaces old estimates
epsilon = 0.2 # exploration rate: random action with prob ε, else best-known
Q = {} # Q table: {(state, action): value}
def act(state):
if random.random() < epsilon: # explore: random move
return random.randint(0, 3)
vals = [Q.get((state, a), 0.0) for a in range(4)] # exploit: pick the max
return vals.index(max(vals))
def step(state, action):
# simulate the environment's transition (in real projects this is env.step)
next_state, reward, done = simulate(state, action)
return next_state, reward, done
for episode in range(5000):
state = 0 # always start from the start cell
done = False
while not done:
a = act(state)
next_state, r, done = step(state, a)
old = Q.get((state, a), 0.0)
# the core update: Q-learning's Bellman update
# new value = immediate reward + γ·(best value at the next state)
best_next = max(Q.get((next_state, a2), 0.0) for a2 in range(4))
Q[(state, a)] = old + alpha * (r + gamma * best_next - old)
state = next_stateTake these dozen-odd lines apart and you'll find they already contain every core ingredient of RL:
| Line of code | The RL concept behind it |
|---|---|
gamma = 0.9 | Discount factor: decay of distant rewards |
epsilon = 0.2 | The exploration–exploitation trade-off (ε-greedy); see Multi-Armed Bandits and Exploration vs Exploitation |
Q.get((state, a)) | The action-value function Q(s,a): scoring "how good is doing this, here" |
act() taking the max Q | Deriving a policy from values: acting greedily |
r + gamma * best_next - old | The TD error (temporal-difference error): the gap between prediction and "reality + expectation" |
Q = Q + alpha * (TD error) | The basic update of value learning: nudge the estimate toward the TD error |
After 5,000 episodes, the path from the start reliably converges to S → right → right → down → down → G, dodging the trap.
Three transferable takeaways from this example
- The policy is never learned directly: here we only learn a Q table; the policy is "take the action with the highest Q in each state" — that's the value-learning paradigm.
- The update uses only "one step of reality + one estimate": this is the heart of TD learning, with far lower variance than waiting for a full episode to finish before updating (Monte Carlo).
- Swap the table for a network and you get DQN: replace
Q[(s,a)]with a neural networkQ(s,a;θ), add experience replay and a target network, and you have the entire starting point of DQN on the Value Learning page.
A fully runnable version (three CartPole implementations) lives in the Progressive Gymnasium Tutorial. Budget two hours to get it running — worth more than reading this page three times.
6. A Boundary Table: What RL Is (and Isn't) Good For
RL is not a universal hammer. Before asking "should we use RL?" in an engineering meeting, check this boundary table first:
| Situation | Is RL a good fit? | Why, and what to use instead |
|---|---|---|
| You need sequential decision-making, where actions affect future states | ✅ Great fit | RL's home turf: games, control, dialogue, ranking |
| The task reduces to "one choice, immediate feedback" | ⚠️ Overkill | A contextual bandit is enough — see Multi-Armed Bandits |
| You have plenty of supervised labels — you just lack a teacher for the policy | ⚠️ Try imitation learning first | Behavior cloning / imitation learning; see RL vs Neighboring Paradigms |
| The environment model is known precisely and the problem size is manageable | ⚠️ Prefer classical methods | Dynamic programming, optimal control (LQR/MPC); see Model-Based RL |
| Every interaction is extremely expensive (surgery, real money) | ⚠️ Proceed with caution | Consider Offline RL, or simulate first |
| You need provably correct behavior (e.g., traffic-signal logic) | ❌ Poor fit | Rules and heuristics are more interpretable and verifiable |
| The goal is hard to write as a scalar reward | ❌ Very hard | Do Reward Engineering first, or the model will inevitably farm the reward |
| Agent actions could cause dangerous consequences | ⚠️ Guardrails required | Add safety constraints and human fallback; see Anatomy of an RL System |
The costliest misuse
Forcing RL onto a problem where interactions are extremely expensive, the environment can't be simulated, and trial-and-error can't be afforded — while the reward is also misspecified. That's three landmines at once. The only sane path here: start with offline data plus behavior cloning as a baseline, then transition gradually; see the decision tree in Offline RL.
7. A Guide to the Rest of the Site
By now you have a foundation-level understanding of RL. Where to go next depends on your goal:
- Want the mathematical framework, systematically: read Markov Decision Process (the five-tuple, Bellman equations), with the Math Primer as a companion.
- Want to know how values are learned: read Value Learning (the full thread from dynamic programming to DQN).
- Want to learn policies directly: read Policy Gradient Methods (REINFORCE → PPO).
- Want an "epic" example first, to build faith: read AlphaGo and Monte Carlo Tree Search and watch "learning × search" beat the best humans.
- Want to get your hands dirty: jump straight into the Progressive Gymnasium Tutorial, or pick a route first via Learning Paths: Three Routes.
- Want the boundaries against other fields nailed down: head to RL vs Neighboring Paradigms.
One last word: RL is a subject you "learn fast and forget fast" — concepts leak away if you don't see them twice. The good news: every concept page on this site ships with a four-piece kit of "intuition + formula + example + pitfalls," so coming back for review costs little.
Further Reading
- RL vs Neighboring Paradigms — expands this page's "how RL differs" sketch into a full seven-paradigm analysis.
- Markov Decision Process — the rigorous mathematics of the agent–environment loop: the five-tuple, return, and Bellman equations.
- Multi-Armed Bandits — RL with the time dimension removed: the minimal laboratory for exploration and exploitation.
- Value Learning — the complete upgrade path for this page's Q-learning example: DP → MC/TD → the DQN family.
- Policy Gradient Methods — the second main line, running parallel to value learning: learning "what to do" directly.
- AlphaGo and Monte Carlo Tree Search — understand the ceiling of "RL + search" through a real historical case.
References
- Sutton, R. S. & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. Full text online: http://incompleteideas.net/book/the-book-2nd.html — Chapter 1 (the reward hypothesis) and Chapter 3 (the agent–environment interface) are the direct sources of this page.
- OpenAI (2018). Spinning Up in Deep RL — Key Concepts. https://spinningup.openai.com/en/latest/spinningup/rl_intro.html — another crisp formulation of "what RL is," including a discussion of the reward hypothesis.
- David Silver (2015). UCL Course on Reinforcement Learning, Lecture 1: Introduction to Reinforcement Learning. https://www.davidsilver.uk/teaching/ — the classic lecture slides where the agent–environment loop comes from.
- Farama Foundation. Gymnasium Documentation. https://gymnasium.farama.org/ — the official entry point for runnable versions of the grid world and CartPole.
- Watkins, C. J. C. H. & Dayan, P. (1992). Q-learning. Machine Learning, 8(3-4), 279–292. https://link.springer.com/article/10.1007/BF00992698 — the original source of the Q-learning example in Section 5.