Skip to content

What Is Reinforcement Learning

On this page Reinforcement learning is the paradigm of learning "what to do" through trial and error — no labels, only after-the-fact rewards. This page gives a precise definition, the agent–environment interaction framework, the reward hypothesis, a three-way comparison with supervised and unsupervised learning, and a complete walkthrough of a minimal RL example.

What Is Reinforcement Learning ​

One-line pitch: this page pins down exactly what the words "reinforcement learning" mean — what problem it solves, what interaction framework it uses, how it differs from the supervised and unsupervised learning you already know, and what a minimal example looks like. By the end, you'll be able to tell anyone: RL is the machine learning paradigm where an agent learns what to do by trial and error in an environment, guided by rewards that arrive after the fact.

1. A One-Sentence Definition and the Agent–Environment Loop ​

Reinforcement Learning (RL) studies how an agent learns a policy through repeated interaction with an environment, so as to maximize the cumulative reward accumulated over the long run.

The catch comes down to three things RL does not have:

  • No labels: nobody tells you "the correct action right now is a*";
  • No immediate feedback: the consequences of an action — good or bad — may only show up much later;
  • No independent and identically distributed data: the state you see next depends on what you just did — the data is generated, and contaminated, by your own behavior.

That is the deepest difference from supervised learning: supervised learning learns "what is this"; RL learns "what should I do".

The interaction framework fits in one diagram:

text
                ┌────────────────────────────┐
                │        Environment         │
                │  (rules, simulator,        │
                │   opponents, users)        │
                └────────────────────────────┘
                     ▲               │
        state s_t    │               │  action a_t
       (+ reward r_t)│               ▼
                ┌────────────────────────────┐
                │           Agent            │
                │  (policy π, value function,│
                │   memory)                  │
                └────────────────────────────┘

What happens at each time step t is, mathematically, one sample from a Markov decision process (MDP) — see Markov Decision Process for the full definition. For now, the plain-spoken version:

text
Loop:
1. The environment is in state s_t and hands s_t (plus reward r_t) to the agent
2. The agent picks action a_t according to policy π(a|s)
3. The environment produces the next state s_{t+1} via transition function P(s_{t+1} | s_t, a_t)
4. The environment produces the immediate reward r_{t+1} via reward function R(s_t, a_t)
5. Go back to 1, until a terminal state (or the budget runs out)

If you only read one thing, read this

The loop above is the skeleton of the entire building. Every algorithm that follows — Q-learning, DQN, PPO, SAC, RLHF — is a different answer to the same question: "how do I update the agent's internal parameters inside this loop?" Burn this diagram into your brain; when you later read Anatomy of an RL System, you'll watch it scale up into a six-layer engineering system.

2. The Reward Hypothesis and the Four Elements ​

1. The Reward Hypothesis ​

RL rests on an axiom known as the reward hypothesis, stated explicitly by Sutton & Barto:

Any goal can be formulated as the maximization of the expected value of a scalar reward signal (cumulative reward).

"Live a long, healthy life" becomes "+1 for every extra day alive." "Win a game of Go" becomes "+1 for a win, −1 for a loss." "Recommend content users like" becomes "+0.1 per click, +0.5 per purchase." Whether this translation is any good decides whether your RL project lives or dies — which is the entire reason the Reward Engineering page exists.

The reward hypothesis is a hypothesis

Translating a goal into a reward is almost always lossy: get the translation wrong, and the agent will game the reward instead of doing what you want (reward hacking). The classic example: a cleaning robot discovers it earns more cleaning reward by dirtying the rug and then cleaning it. This isn't a bug — it's a reward design problem. See the case collection in Reward Engineering.

2. The Four Elements: State, Action, Policy, Return ​

ElementEnglishDefinitionAn intuition (CartPole)
StateState sEverything the environment exposes to the agent at a given momentCart position and velocity, pole angle and angular velocity
ActionAction aOne choice the agent can makePush left, push right
PolicyPolicy π(a|s)A mapping from states to actions (or distributions over actions)Pole tilts far left → push left
ReturnReturn G_tThe (discounted) sum of rewards from time t onwardThe total of "not falling" rewards in the future

The definition of return (with discount factor γ):

text
G_t = r_{t+1} + γ·r_{t+2} + γ²·r_{t+3} + ... = Σ_{k=0}^{∞} γ^k · r_{t+k+1}

The discount factor γ ∈ [0, 1) does three jobs:

  1. Mathematically, it makes the infinite sum converge;
  2. It encodes "a dollar today is worth more than a dollar tomorrow" — decision-makers naturally prefer earlier rewards;
  3. It controls the agent's "foresight": the closer γ is to 1, the more long-sighted the agent; the smaller γ, the more myopic.

Why γ ≠ 1 is the sensible choice

If γ = 1 and the task never ends, the return can diverge; even when it doesn't, "one dollar today equals one dollar ten years from now" defies common sense. In most engineering settings, γ sits between 0.9 and 0.99. γ is the first hyperparameter you'll touch in Tuning in Practice.

3. A Three-Way Comparison with Supervised and Unsupervised Learning ​

RL is often called the "third paradigm" of machine learning. Here's a side-by-side look along three dimensions:

DimensionSupervised learningUnsupervised learningReinforcement learning
Form of dataInput–label pairs (x, y)Inputs x onlyState–action–reward trajectories (s, a, r) sequences
Feedback signalImmediate, explicit label for every sampleNo labels; relies on intrinsic structure of dataAfter-the-fact, sparse, delayed scalar rewards
Learning goalFit an input→output mappingDiscover distributions/structure (clustering, compression)Maximize long-term cumulative return through interaction
Who produces the dataGiven externally (a fixed dataset)Given externallyThe agent's own behavior generates the follow-up data
Can errors be corrected immediately?Yes — the loss hands you gradients directlyNo clear right or wrongNo — an action may have to "take the blame" for outcomes long afterward
Typical tasksClassification, regression, detectionClustering, dimensionality reduction, generationGames, control, dialogue, ranking, alignment
Success metricTest-set accuracyCluster quality / reconstruction errorTrajectory return, task success rate, sample efficiency

The key difference boils down to one sentence: supervised learning's data is stationary; RL's data is dynamic — you are changing the very distribution you will learn from. This distribution shift (non-stationarity) is the root of all RL difficulty, and it explains why RL needs exploration and why offline data can serve as a safety net.

Another angle on the same idea

Think of supervised learning as a course where the teacher grades every homework assignment, and RL as a course where you only get a final grade and nothing is ever graded in between — and where each assignment you submit changes what the next class covers. To sharpen the boundaries further (RL vs optimal control, RL vs behavior cloning), head straight to RL vs Neighboring Paradigms.

4. Why "Sequences of Decisions" Are Hard: Credit Assignment and Delayed Rewards ​

A single decision is easy: evaluate the immediate reward of each option and pick the largest. RL, however, deals with sequential decision-making — an action you take today may pay off a thousand steps later. This raises two fundamental problems.

1. The Credit Assignment Problem ​

When the final reward finally arrives, which of your past actions should get the credit?

text
Scene: after 10 moves of chess, move 3 plants a seed,
       and move 10 wins the game.
Question: who gets the +1? Move 3? Or move 10?

Credit only the last move, and the agent never learns to "set things up"; spread it evenly across every move, and move 3's contribution gets diluted. TD learning, eligibility traces, GAE, value networks — half of all RL algorithms are answers to this very question.

2. Delayed and Sparse Rewards ​

Many real-world tasks offer extremely sparse rewards: a robot may take ten thousand steps before hitting a single reward (grasping an object, reaching a goal). With rewards that rare, random exploration almost never stumbles onto one, and learning stalls. Remedies include reward shaping, curriculum learning, and intrinsic rewards (curiosity) — see Exploration vs Exploitation and Reward Engineering for details.

A common beginner's mistake

Assuming that "designing a dense reward will fix it." A dense reward does ease sparsity, but it brings reward hacking and the risk of "locally optimal reward-farmers" (a robot that learns to spin in place farming reward, for instance). Reward density is not the higher the better — it's a trade-off between density and resistance to being gamed.

3. Why Look-Up Tables Fail in the Real World ​

Credit assignment + sparse rewards + high-dimensional states: three mountains that crush any "table lookup" method. Real state spaces are continuous and vast (images, sensor readings, text) — you cannot store a value per state. The fix is function approximation: use a neural network to learn the mapping from "state → value/action distribution." That is exactly what the two main lines — Value Learning and Policy Gradient Methods — each set out to do.

5. A Minimal RL Example: Q-Learning in a Grid World ​

Enough theory — let's dissect a minimal example end to end: a 3×3 grid world.

text
┌────┬────┬────┐
│ S  │ ·  │ ·  │    S = start
├────┼────┼────┤    G = goal (reward +10, terminate)
│ ·  │ ✗  │ ·  │    ✗ = trap (reward -5, terminate)
├────┼────┼────┤    other cells: -0.1 per step (hurry up!)
│ ·  │ ·  │ G  │    actions: up/down/left/right
└────┴────┴────┘

The agent doesn't know the map; it can only learn by "taking a step and seeing the reward." We solve it with Q-learning — a Q table storing "the long-term value of doing each action in each cell":

python
import random

# states: 9 cells numbered 0~8; actions: 0=up 1=down 2=left 3=right
gamma = 0.9        # discount factor: closer to 1 = more foresight
alpha = 0.1        # learning rate: how fast new info replaces old estimates
epsilon = 0.2      # exploration rate: random action with prob ε, else best-known

Q = {}             # Q table: {(state, action): value}

def act(state):
    if random.random() < epsilon:        # explore: random move
        return random.randint(0, 3)
    vals = [Q.get((state, a), 0.0) for a in range(4)]  # exploit: pick the max
    return vals.index(max(vals))

def step(state, action):
    # simulate the environment's transition (in real projects this is env.step)
    next_state, reward, done = simulate(state, action)
    return next_state, reward, done

for episode in range(5000):
    state = 0                            # always start from the start cell
    done = False
    while not done:
        a = act(state)
        next_state, r, done = step(state, a)
        old = Q.get((state, a), 0.0)
        # the core update: Q-learning's Bellman update
        # new value = immediate reward + γ·(best value at the next state)
        best_next = max(Q.get((next_state, a2), 0.0) for a2 in range(4))
        Q[(state, a)] = old + alpha * (r + gamma * best_next - old)
        state = next_state

Take these dozen-odd lines apart and you'll find they already contain every core ingredient of RL:

Line of codeThe RL concept behind it
gamma = 0.9Discount factor: decay of distant rewards
epsilon = 0.2The exploration–exploitation trade-off (ε-greedy); see Multi-Armed Bandits and Exploration vs Exploitation
Q.get((state, a))The action-value function Q(s,a): scoring "how good is doing this, here"
act() taking the max QDeriving a policy from values: acting greedily
r + gamma * best_next - oldThe TD error (temporal-difference error): the gap between prediction and "reality + expectation"
Q = Q + alpha * (TD error)The basic update of value learning: nudge the estimate toward the TD error

After 5,000 episodes, the path from the start reliably converges to S → right → right → down → down → G, dodging the trap.

Three transferable takeaways from this example

  1. The policy is never learned directly: here we only learn a Q table; the policy is "take the action with the highest Q in each state" — that's the value-learning paradigm.
  2. The update uses only "one step of reality + one estimate": this is the heart of TD learning, with far lower variance than waiting for a full episode to finish before updating (Monte Carlo).
  3. Swap the table for a network and you get DQN: replace Q[(s,a)] with a neural network Q(s,a;θ), add experience replay and a target network, and you have the entire starting point of DQN on the Value Learning page.

A fully runnable version (three CartPole implementations) lives in the Progressive Gymnasium Tutorial. Budget two hours to get it running — worth more than reading this page three times.

6. A Boundary Table: What RL Is (and Isn't) Good For ​

RL is not a universal hammer. Before asking "should we use RL?" in an engineering meeting, check this boundary table first:

SituationIs RL a good fit?Why, and what to use instead
You need sequential decision-making, where actions affect future states✅ Great fitRL's home turf: games, control, dialogue, ranking
The task reduces to "one choice, immediate feedback"⚠️ OverkillA contextual bandit is enough — see Multi-Armed Bandits
You have plenty of supervised labels — you just lack a teacher for the policy⚠️ Try imitation learning firstBehavior cloning / imitation learning; see RL vs Neighboring Paradigms
The environment model is known precisely and the problem size is manageable⚠️ Prefer classical methodsDynamic programming, optimal control (LQR/MPC); see Model-Based RL
Every interaction is extremely expensive (surgery, real money)⚠️ Proceed with cautionConsider Offline RL, or simulate first
You need provably correct behavior (e.g., traffic-signal logic)❌ Poor fitRules and heuristics are more interpretable and verifiable
The goal is hard to write as a scalar reward❌ Very hardDo Reward Engineering first, or the model will inevitably farm the reward
Agent actions could cause dangerous consequences⚠️ Guardrails requiredAdd safety constraints and human fallback; see Anatomy of an RL System

The costliest misuse

Forcing RL onto a problem where interactions are extremely expensive, the environment can't be simulated, and trial-and-error can't be afforded — while the reward is also misspecified. That's three landmines at once. The only sane path here: start with offline data plus behavior cloning as a baseline, then transition gradually; see the decision tree in Offline RL.

7. A Guide to the Rest of the Site ​

By now you have a foundation-level understanding of RL. Where to go next depends on your goal:

One last word: RL is a subject you "learn fast and forget fast" — concepts leak away if you don't see them twice. The good news: every concept page on this site ships with a four-piece kit of "intuition + formula + example + pitfalls," so coming back for review costs little.

Further Reading ​

References ​