Theme
Reinforcement Learning
Concept Definition: Learning to Decide through Trial and Error
Supervised learning has labels, unsupervised learning finds structure, and reinforcement learning (RL) faces a stricter setting: no answer key, only "delayed rewards." An agent interacts with an environment: observes state s, takes action a, the environment returns a new state s' and a reward r, and the cycle repeats. The goal of RL is to learn a policy (π: state → action) that maximizes long-term cumulative reward.
┌─────────────── Agent ───────────────┐
│ │
│ Policy π(a|s) ←── Learning Target │
└──┬──────────────┬───────────────────┘
│ Action a │ Observation s, Reward r
▼ ▲
┌───────────────────────────────────┐
│ Environment (dynamics) │
└───────────────────────────────────┘The essential difference from supervised/unsupervised:
| Dimension | Supervised Learning | Reinforcement Learning |
|---|---|---|
| Data form | Labeled (x, y) | Trajectories (s, a, r, s') from interaction |
| Feedback | Immediate, every sample | Delayed, sparse (reward may come dozens of steps later) |
| Goal | Fit a mapping | Optimize long-term cumulative return |
| Key challenge | Generalization | Credit assignment (which step is responsible for the final outcome) + exploration |
The credit assignment problem is the core challenge of RL: a chess game was lost, but was it the mistake at move 10 or move 40? When reward only appears at the end, how to "assign" credit/responsibility to each intermediate step — this is the purpose of value functions.
Formal Framework: Markov Decision Process (MDP)
Almost all RL problems can be formulated as an MDP, with five elements:
MDP = (S, A, P, R, γ)
S Set of states A Set of actions
P State transition prob R Reward function
γ Discount factor (0≤γ<1): rewards further in the future are discounted moreMarkov property: the next state depends only on the current state and action, not on history (p(s'|s,a)). The discount factor γ expresses "a dollar today is worth more than a dollar tomorrow" — the closer γ is to 1, the more "far-sighted"; the smaller γ, the more "short-sighted."
The goal of RL is to find a policy π that maximizes cumulative discounted return G = Σ γᵗ rₜ.
Exploration vs. Exploitation: RL's Primary Contradiction
Exploitation: act according to the currently known best strategy (safe); Exploration: try unknown actions, potentially discovering a better strategy (risky). The two are in inherent conflict — the multi-armed bandit is a simplified model for studying this problem.
| Strategy | Idea |
|---|---|
| ε-greedy | Explore randomly with probability ε, otherwise greedily exploit; ε decays over time |
| UCB | Prioritize actions with "high mean + high uncertainty" (upper confidence bound) |
| Thompson Sampling | Sample actions by posterior probability, Bayesian perspective |
A practical rule of thumb: explore more early in training (large ε), exploit more later (small ε). In real business (recommendation, advertising), exploration is often designed as a separate system component because exploration directly costs short-term revenue.
Two Major Learning Paradigms
1. Value Learning: Learning "How Good is This State"
- State value V(s): the expected return starting from state s and following policy π;
- Action value Q(s,a): the expected return after taking action a in state s.
Q-learning (classic algorithm): iteratively updates the Q-table using the Bellman equation:
Q(s,a) ← Q(s,a) + α [ r + γ·maxₐ' Q(s',a') − Q(s,a) ]Where r + γ·max Q(s',a') is the "TD target" (current reward + optimal value of next state), and the term in parentheses is the TD error — using a "predicted next step" to correct "the current estimate."
DQN (Deep Q-Network, 2015): uses a neural network to approximate the Q function, solving for high-dimensional states (e.g., game screens). Two key engineering tricks:
- Experience Replay: store historical transitions (s,a,r,s') in a buffer and sample randomly for training — breaks sample correlation;
- Target Network: the TD target is computed by an independent network that updates slowly, stabilizing training.
2. Policy Learning: Directly Learning "What to Do"
Policy Gradient: directly optimize the policy network π(a|s), in the gradient direction = "increase the probability of high-return actions." PPO (Proximal Policy Optimization, 2017) is the de facto standard in the policy gradient family: it limits the magnitude of policy updates per step by clipping the objective function, making training stable and implementation simple. It is the default choice for modern RL applications (games, robotics, large model alignment).
Actor-Critic: Combining Value + Policy
Actor (policy network, decides actions) + Critic (value network, evaluates how good an action is). The Critic provides a more fine-grained signal than the raw reward (the advantage function), and the Actor uses this signal to update the policy. A3C, PPO, SAC all belong to the Actor-Critic family — it alleviates both the high variance of value learning and the slow convergence of policy learning.
Classic Achievements of Deep Reinforcement Learning
| Year | Achievement | Significance |
|---|---|---|
| 2013/2015 | DQN outperforming humans on Atari | Beginning of deep learning + RL (Nature cover) |
| 2016 | AlphaGo defeats Lee Sedol 4:1 | Value network + policy network + Monte Carlo tree search (MCTS), AI goes mainstream |
| 2017 | AlphaZero self-learns Go/chess/shogi from scratch | No human game records needed, pure self-play reinforcement learning |
| 2019 | OpenAI Five / AlphaStar defeats pro players in Dota 2 / StarCraft | Complex multi-agent scenarios |
| 2022+ | ChatGPT's RLHF | RL enters large models: using human feedback as reward, aligning model behavior |
The last row deserves elaboration: RLHF (Reinforcement Learning from Human Feedback) turns "human preferences" into a reward model, then uses PPO to optimize the language model — this is the key technology that takes large models "from being able to talk to being able to chat," and is also the first time RL became a core component of a general-purpose product. See Large Language Models (LLM).
What Problems is RL Suited For?
RL isn't a silver bullet; it has clear applicability boundaries:
| Suitable | Not Suitable |
|---|---|
| Decisions are sequential (one step affects the next) | Single-step decisions (use supervised learning/rules) |
| Low-cost trial and error is possible (simulators, games, log replay) | Trial-and-error is extremely costly (surgery, lending) |
| Reward is definable (win/loss, conversion, energy consumption) | Reward can't be quantified or is easily gamed |
| Environment can be simulated or has lots of interaction data | Data is only passively observable (offline RL still hard) |
Engineering reality: the biggest cost of RL projects is the environment (simulator). AlphaGo succeeded because Go has a perfect simulator; robot RL is hard to land because real physical environments are too expensive to explore. The vast majority of enterprise RL applications focus on recommendation, advertising, and scheduling — scenarios where "offline learning from logs + small-step exploration online" is feasible.
Reward design is the devil in RL
Get the reward function wrong, and the agent will "game the score": give a robot a task to clean up trash, and it might first make trash then clean it (more reward); give a model the task of writing ad copy, and it might pile on "attention-grabbing" words. Reward hacking is the number one trap in RL engineering. The countermeasure is to combine reward design with safety guardrails — especially critical in large model RLHF (where the model learns to "please the annotator" rather than "tell the truth").
Tradeoffs
- Value learning vs. Policy learning: continuous action space (robots, control) → policy gradient / PPO (Q-learning doesn't apply to continuous actions); discrete actions → either works.
- Online vs. Offline RL: can interact → use online (PPO/SAC); only historical data → use offline RL (CQL, etc.), but offline RL is sensitive to data distribution and hard to land.
- Model-based vs. Model-free: with environment dynamics model (model-based) → high sample efficiency; model-free → general-purpose but low sample efficiency. Model-based yields big returns in games/simulations.
- Exploration cost: too little exploration means missing the optimal; too much wastes resources — ε, UCB parameters need tuning per scenario.
Further Reading
- Supervised Learning — comparison of the three paradigms
- Unsupervised Learning — the second paradigm
- RL Applications — RL in AlphaGo, robotics, recommendation
- Large Language Models (LLM) — RLHF and alignment
- Math Primer — dynamic programming behind the Bellman equation
References
- Sutton & Barto. Reinforcement Learning: An Introduction (2nd ed., 2018) — the RL bible, available free online
- Mnih et al. Human-level control through deep reinforcement learning (DQN, Nature 2015)
- Schulman et al. Proximal Policy Optimization Algorithms (PPO, 2017)
- Silver et al. Mastering the game of Go with deep neural networks and tree search (AlphaGo, Nature 2016)
- Ouyang et al. Training language models to follow instructions with human feedback (InstructGPT/RLHF, 2022)
- OpenAI Spinning Up in Deep RL tutorial — the best resource for getting started in deep RL