Skip to content

Reinforcement Learning

Quick overview Reinforcement learning is the third modeling paradigm: no labels, no answer key, only reward signals from the environment. This article clarifies the MDP framework, exploration vs. exploitation, value learning vs. policy learning, deep reinforcement learning (DQN/PPO), and what problems RL is suited for.

Reinforcement Learning ​

Concept Definition: Learning to Decide through Trial and Error ​

Supervised learning has labels, unsupervised learning finds structure, and reinforcement learning (RL) faces a stricter setting: no answer key, only "delayed rewards." An agent interacts with an environment: observes state s, takes action a, the environment returns a new state s' and a reward r, and the cycle repeats. The goal of RL is to learn a policy (π: state → action) that maximizes long-term cumulative reward.

          ┌─────────────── Agent ───────────────┐
          │                                      │
          │  Policy π(a|s)  ←── Learning Target  │
          └──┬──────────────┬───────────────────┘
             │ Action a      │ Observation s, Reward r
             ▼              ▲
          ┌───────────────────────────────────┐
          │           Environment (dynamics)   │
          └───────────────────────────────────┘

The essential difference from supervised/unsupervised:

DimensionSupervised LearningReinforcement Learning
Data formLabeled (x, y)Trajectories (s, a, r, s') from interaction
FeedbackImmediate, every sampleDelayed, sparse (reward may come dozens of steps later)
GoalFit a mappingOptimize long-term cumulative return
Key challengeGeneralizationCredit assignment (which step is responsible for the final outcome) + exploration

The credit assignment problem is the core challenge of RL: a chess game was lost, but was it the mistake at move 10 or move 40? When reward only appears at the end, how to "assign" credit/responsibility to each intermediate step — this is the purpose of value functions.

Formal Framework: Markov Decision Process (MDP) ​

Almost all RL problems can be formulated as an MDP, with five elements:

MDP = (S, A, P, R, γ)
 S  Set of states          A  Set of actions
 P  State transition prob  R  Reward function
 γ  Discount factor (0≤γ<1): rewards further in the future are discounted more

Markov property: the next state depends only on the current state and action, not on history (p(s'|s,a)). The discount factor γ expresses "a dollar today is worth more than a dollar tomorrow" — the closer γ is to 1, the more "far-sighted"; the smaller γ, the more "short-sighted."

The goal of RL is to find a policy π that maximizes cumulative discounted return G = Σ γᵗ rₜ.

Exploration vs. Exploitation: RL's Primary Contradiction ​

Exploitation: act according to the currently known best strategy (safe); Exploration: try unknown actions, potentially discovering a better strategy (risky). The two are in inherent conflict — the multi-armed bandit is a simplified model for studying this problem.

StrategyIdea
ε-greedyExplore randomly with probability ε, otherwise greedily exploit; ε decays over time
UCBPrioritize actions with "high mean + high uncertainty" (upper confidence bound)
Thompson SamplingSample actions by posterior probability, Bayesian perspective

A practical rule of thumb: explore more early in training (large ε), exploit more later (small ε). In real business (recommendation, advertising), exploration is often designed as a separate system component because exploration directly costs short-term revenue.

Two Major Learning Paradigms ​

1. Value Learning: Learning "How Good is This State" ​

  • State value V(s): the expected return starting from state s and following policy π;
  • Action value Q(s,a): the expected return after taking action a in state s.

Q-learning (classic algorithm): iteratively updates the Q-table using the Bellman equation:

Q(s,a) ← Q(s,a) + α [ r + γ·maxₐ' Q(s',a') − Q(s,a) ]

Where r + γ·max Q(s',a') is the "TD target" (current reward + optimal value of next state), and the term in parentheses is the TD error — using a "predicted next step" to correct "the current estimate."

DQN (Deep Q-Network, 2015): uses a neural network to approximate the Q function, solving for high-dimensional states (e.g., game screens). Two key engineering tricks:

  • Experience Replay: store historical transitions (s,a,r,s') in a buffer and sample randomly for training — breaks sample correlation;
  • Target Network: the TD target is computed by an independent network that updates slowly, stabilizing training.

2. Policy Learning: Directly Learning "What to Do" ​

Policy Gradient: directly optimize the policy network π(a|s), in the gradient direction = "increase the probability of high-return actions." PPO (Proximal Policy Optimization, 2017) is the de facto standard in the policy gradient family: it limits the magnitude of policy updates per step by clipping the objective function, making training stable and implementation simple. It is the default choice for modern RL applications (games, robotics, large model alignment).

Actor-Critic: Combining Value + Policy ​

Actor (policy network, decides actions) + Critic (value network, evaluates how good an action is). The Critic provides a more fine-grained signal than the raw reward (the advantage function), and the Actor uses this signal to update the policy. A3C, PPO, SAC all belong to the Actor-Critic family — it alleviates both the high variance of value learning and the slow convergence of policy learning.

Classic Achievements of Deep Reinforcement Learning ​

YearAchievementSignificance
2013/2015DQN outperforming humans on AtariBeginning of deep learning + RL (Nature cover)
2016AlphaGo defeats Lee Sedol 4:1Value network + policy network + Monte Carlo tree search (MCTS), AI goes mainstream
2017AlphaZero self-learns Go/chess/shogi from scratchNo human game records needed, pure self-play reinforcement learning
2019OpenAI Five / AlphaStar defeats pro players in Dota 2 / StarCraftComplex multi-agent scenarios
2022+ChatGPT's RLHFRL enters large models: using human feedback as reward, aligning model behavior

The last row deserves elaboration: RLHF (Reinforcement Learning from Human Feedback) turns "human preferences" into a reward model, then uses PPO to optimize the language model — this is the key technology that takes large models "from being able to talk to being able to chat," and is also the first time RL became a core component of a general-purpose product. See Large Language Models (LLM).

What Problems is RL Suited For? ​

RL isn't a silver bullet; it has clear applicability boundaries:

SuitableNot Suitable
Decisions are sequential (one step affects the next)Single-step decisions (use supervised learning/rules)
Low-cost trial and error is possible (simulators, games, log replay)Trial-and-error is extremely costly (surgery, lending)
Reward is definable (win/loss, conversion, energy consumption)Reward can't be quantified or is easily gamed
Environment can be simulated or has lots of interaction dataData is only passively observable (offline RL still hard)

Engineering reality: the biggest cost of RL projects is the environment (simulator). AlphaGo succeeded because Go has a perfect simulator; robot RL is hard to land because real physical environments are too expensive to explore. The vast majority of enterprise RL applications focus on recommendation, advertising, and scheduling — scenarios where "offline learning from logs + small-step exploration online" is feasible.

Reward design is the devil in RL

Get the reward function wrong, and the agent will "game the score": give a robot a task to clean up trash, and it might first make trash then clean it (more reward); give a model the task of writing ad copy, and it might pile on "attention-grabbing" words. Reward hacking is the number one trap in RL engineering. The countermeasure is to combine reward design with safety guardrails — especially critical in large model RLHF (where the model learns to "please the annotator" rather than "tell the truth").

Tradeoffs ​

  • Value learning vs. Policy learning: continuous action space (robots, control) → policy gradient / PPO (Q-learning doesn't apply to continuous actions); discrete actions → either works.
  • Online vs. Offline RL: can interact → use online (PPO/SAC); only historical data → use offline RL (CQL, etc.), but offline RL is sensitive to data distribution and hard to land.
  • Model-based vs. Model-free: with environment dynamics model (model-based) → high sample efficiency; model-free → general-purpose but low sample efficiency. Model-based yields big returns in games/simulations.
  • Exploration cost: too little exploration means missing the optimal; too much wastes resources — ε, UCB parameters need tuning per scenario.

Further Reading ​

References ​