Appearance
A Brief History of Reinforcement Learning
In one sentence: this page compresses seventy years of RL into a single traceable timeline — from Bellman's equation to the RLHF pipeline. You'll watch six tributaries (dynamic programming, tabular methods, policy gradients, deep RL, multi-agent RL, and RLHF) rise, branch apart, and finally converge in today's deep RL. When you finish, you'll be able to say what problem each key concept was born to answer — the first step toward reading the paper map.
1. Overview: Seventy Years, Three Eras
Slice the history of RL from 1950 to 2025 into three eras, each of which solved one fundamental problem:
text
Era 1 (1950s–1980s) Foundations: "learning is possible," proved mathematically
· Bellman equation, MDP formalization
· The seeds of TD thinking; the credit assignment problem posed
Era 2 (1980s–2013) Tabular methods mature: "learning works" in finite state spaces
· TD(λ), Q-learning, TD-Gammon
· Convergence proven mathematically, but scale stayed out of reach
Era 3 (2013–present) Deep RL: "learning scales" to high-dimensional states
· The DQN/PPO/AlphaGo trinity
· Scaling up: RLHF, world models, scalable RLWhat bridges each era is representational power: tabular methods were stuck on the size of the state space, and deep learning swept RL into vision, language, and physical control in a single move. Let's walk through it decade by decade.
2. 1950s: Bellman's Dynamic Programming and the MDP Formalization
It all starts with two developments:
- Markov chain theory (A. A. Markov, from 1906): a stochastic process in which "the next step depends only on the current state" — the groundwork for the Markov property, the idea that the state fully captures the past.
- Richard Bellman invented dynamic programming in the 1950s: decompose "multi-step optimal decision-making" into "the current step plus the optimal solution of what remains," yielding the famous Bellman equation. His book Dynamic Programming appeared in 1957. In 1960 he went further, proposing the Markov decision process (MDP) as a complete formal framework — the direct ancestor of the five-tuple on today's Markov Decision Process page.
text
Bellman's principle of optimality (in one sentence):
For a policy to be optimal overall, every remaining part of it —
from any moment on — must also be optimal.
→ So V*(s) can be defined recursively: V*(s) = max_a [ R(s,a) + γ·E[V*(s')] ]
→ This recursion is the starting point of every value-learning algorithm today.Why this history matters
The Bellman equation is no "textbook formula" — it was invented in the 1950s as a general principle for sequential decision-making in aerospace control and inventory management. Understand the context of its birth and you understand why RL, operations research, and optimal control share the same roots: all three are built on the same equation. For a paper-level narrative, see the Bellman chapter in Classic Papers, Closely Read.
3. 1960s–1980s: Trial-and-Error Learning and the Seeds of TD
At this stage "reinforcement learning" wasn't even called reinforcement learning, but three lines of work were quietly taking shape:
1. Trial-and-error learning enters the machine
- Arthur Samuel (1959): his checkers program pioneered "a machine that learns by playing against itself" — the ancestor of self-play, sixty years before AlphaGo.
- Bernard Widrow and Ted Hoff (1960): proposed the LMS learning rule; Widrow later used "add/subtract" reinforcement signals to balance an inverted pendulum — already a sketch of reward-driven adaptive control.
- Marvin Minsky (1961): in Steps Toward Artificial Intelligence, he formally posed the credit assignment problem: "When an entire sequence of actions ends in failure, which step is to blame?" — a question that would govern algorithm design for decades.
2. The formalization of "trial and error"
During the 1960s–70s, psychology experiments — classical conditioning in cuckoos, reward-learning studies — inspired a family of "trial-and-error learners." Harry Klopf (1972) and others proposed the "hedonistic neuron" hypothesis; in the late 1970s, Barto, Sutton, and Anderson (1983) demonstrated trial-and-error learning on the inverted-pendulum task with their "associative search element" (ASE/ACE) networks — universally recognized as the direct forerunner of modern RL algorithms.
Why this "germination period" matters
The contribution of the 1960s–80s wasn't algorithmic strength — it was putting three questions formally on the table: what to learn (policy/value), how to feed back (reward), and when to settle the accounts (credit assignment). The answers took another decade to crystallize into algorithms.
4. 1988–1992: Temporal-Difference Learning and Q-learning
These were the "revolutionary years" of tabular RL, in which two landmarks appeared:
1. Sutton's temporal-difference learning (TD, 1988)
Learning to Predict by the Methods of Temporal Differences — Sutton proved the viability of a "predict-and-correct-as-you-go" method: instead of waiting for the final outcome, update this step's prediction with the next step's prediction:
text
The intuition behind a TD update:
old prediction ──plus──▶ (actual reward + γ·new prediction − old prediction) × learning rate
↑____________ this correction term is the TD error δ_t ____________↑
Monte Carlo: you only find out who's right after the whole episode (high variance, unbiased)
TD: every step corrects a prediction with another prediction (biased, low variance, online)The elegance of TD is that it's incremental: no need to store entire trajectories. This idea later carried half of Value-Based Learning.
2. Watkins's Q-learning (1989 thesis / 1992 formal publication)
Chris Watkins proposed Q-learning in his PhD thesis: rather than learning the state value V, learn the state-action value Q(s,a), and in the update use the Q-value of the optimal next action — which makes Q-learning off-policy (the policy being learned differs from the behavior policy): you can "explore freely while learning the optimal policy." This property was revolutionary, and DQN inherited it directly.
In the same period came Williams's REINFORCE (1992) — the source of policy gradients: directly raise the probability of actions that yield high returns. It defined the second main line, parallel to value learning — policy gradient methods.
Why Q-learning was a "revolution"
Before Q-learning, the mainstream view held that policy and value had to be learned jointly, and on-policy at that. Q-learning proved you can collect data with any exploration policy while simultaneously learning an optimal one — the logical seed for the "experience replay" and "offline RL" that followed. The convergence proof was completed by Watkins & Dayan (1992).
5. 1992–2013: TD-Gammon and the Golden Age of Classic RL
1. Tesauro's TD-Gammon (1992–1995)
Gerald Tesauro trained a backgammon program with TD learning: a neural network served as the value function approximator, trained on nothing but self-play data. The result stunned the backgammon world: TD-Gammon reached the level of top human grandmasters and developed endgame strategies no human had ever played.
TD-Gammon's historical significance lies not in "winning" but in proving two things:
- "Value function approximation + self-play" can beat humans without any human knowledge;
- Neural networks as function approximators — rather than lookup tables — are viable in RL.
It is the direct intellectual ancestor of AlphaGo. TD-Gammon also marks the "golden age of classic RL": from the late 1990s through the 2000s, RL began to flower in robotics, games, and resource scheduling. Kaelbling, Littman, and Moore (1996) published the famous Reinforcement Learning: A Survey, organizing the whole field into a system.
2. Other main lines of the golden age
| Year | Work | Contribution |
|---|---|---|
| 1993–94 | Rummery & Niranjan's SARSA | on-policy TD control, complementing Q-learning |
| 1996 | Bertsekas & Tsitsiklis, Neuro-Dynamic Programming | welding neural networks and dynamic programming together in theory |
| 1999 | Sutton, McAllester, Singh's policy gradient theorem | laying the theoretical foundation for policy gradient methods |
| 2000s | RL in robotics (crawling, inverted pendulum) and game-AI applications | RL moves from the lab toward engineering trials |
| 2003 | R-Max, E³ and other exploration theories | sample-complexity theory for exploration–exploitation |
The golden age's ceiling
Tabular methods plus simple function approximation could solve "low-dimensional" problems, but they were utterly powerless against images (high-dimensional pixels) and language — feature engineering became the bottleneck. This ceiling stood until deep learning entered the scene in 2013.
6. 2013–2017: The Deep RL Revolution (DQN, AlphaGo, PPO)
1. DQN: deep learning takes over value learning (2013/2015)
Mnih et al.'s DQN was the first to learn Q-values directly from pixels with a convolutional neural network, surpassing human performance on 49 Atari games. The recipe for success was two engineering tricks (omit either and it fails):
- Experience replay: store past (s,a,r,s′) transitions in a large buffer and sample randomly, breaking data correlations;
- Target network: a "frozen" copy of the network supplies the TD targets, stabilizing bootstrapped updates.
DQN opened the door to "deep RL." For the full mechanics, see Value-Based Learning; for paper details, the DQN chapter of Classic Papers, Closely Read; and for the whole gaming landscape, Atari and Video Games.
2. AlphaGo: search and learning converge (2016)
Silver et al.'s AlphaGo combined a policy network, a value network, and MCTS search to defeat world champion Lee Sedol at Go — a domain then believed to be "at least a decade away." Its three-stage training (supervised policy network → RL self-play refinement → value network) is covered in full on the AlphaGo and Monte Carlo Tree Search page. AlphaGo Zero (2017) then dropped human game records for pure self-play, and AlphaZero (2017) generalized the same recipe to chess and shogi.
3. PPO: making policy gradients "stable and tunable" (2017)
Schulman et al. proposed PPO: a clipped objective limits the size of each update, solving in one stroke both TRPO's computational complexity and the "update too large, policy collapses" problem:
text
PPO's clipped objective (the intuition):
the "edge" of the new policy over the old one: ratio = π_new / π_old
multiply this ratio by the advantage A, then clip it to [1-ε, 1+ε]
→ good actions become more likely, but never by too much in one step
→ stable, simple, first-order optimization only — which is why it became the industry defaultTogether with A2C, DDPG, TD3, and SAC, PPO forms the actor-critic family; for the genealogy and how to choose, see The Actor-Critic Family.
How the deep RL trinity divided the work
DQN proved "deep networks can learn values," AlphaGo proved "deep networks combined with search can reach superhuman performance," and PPO proved "deep policy learning is stable and reproducible in engineering terms." Each opened a successor line of its own — value, search, and policy, respectively.
7. 2017–2020: Scaling Up and Multi-Agent RL
Deep RL entered the era of "massive compute + many agents":
1. The rise of multi-agent RL
| Year | Work | Significance |
|---|---|---|
| 2017 | MADDPG | representative of the centralized-training-with-decentralized-execution (CTDE) paradigm |
| 2018 | QMIX | value factorization: approximate the joint value while decomposing it per agent |
| 2019 | OpenAI Five (Dota 2) | 5v5 self-play reaches professional level |
| 2019 | AlphaStar (StarCraft II) | milestone for two-player imperfect-information games |
| 2021 | MAPPO | proof that PPO-ified methods also excel in multi-agent settings |
The core difficulties of multi-agent RL (non-stationarity, credit assignment across agents, equilibrium concepts) are covered in Multi-Agent Reinforcement Learning.
2. World models and the engineering of algorithms
- World Models (Ha & Schmidhuber, 2018) and Dreamer (Hafner et al., 2020): generative models that "predict the next frame" entered RL, letting agents learn inside a "dream" — the renaissance of the model-based line; see Model-Based RL.
- SAC (2018): the maximum-entropy objective made continuous control both stable and efficient; now the default in robotics.
- MuZero (2020): the world-model version of AlphaZero — it learns even the environment model itself, reaching superhuman play in games with unknown rules from actions and rewards alone.
8. 2020–Present: RLHF, World Models, and Scalable RL
1. RLHF and LLM alignment
In 2020, Stiennon et al.'s Learning to Summarize from Human Feedback established the "human preferences → reward model → PPO optimization" pipeline; in 2022, InstructGPT scaled it to ChatGPT-sized models; in 2023, DPO proved that preferences can be optimized directly, no explicit reward model needed. This line has since grown into alignment as a standalone field — the full pipeline is on RLHF and Alignment from Human Feedback, with the hands-on walkthrough in LLM Alignment: RLHF in Practice.
text
The three stages of RLHF (settled 2020–2022):
SFT (imitate demonstrations) → reward model (learn to score from human preference pairs)
→ PPO (fine-tune the policy against the reward model, anchored to a reference model)
→ the product: a conversational model that actually "speaks human"
The 2023 simplification, DPO: skip the reward model and optimize directly on preference pairs2. The convergence of world models and scalable RL
- DreamerV3 (2023): one algorithm reaching SOTA on 150+ tasks without per-task tuning, proving the "general world-model learner" viable.
- Offline RL matures: CQL, IQL, and friends made "train on historical data only" a reality; see Offline Reinforcement Learning.
- Scalable infrastructure: Brax/PureJaxRL run tens of thousands of parallel environments on a single GPU, cutting the cost of "RL needs massive sampling" by an order of magnitude.
- RL for reasoning (2025): DeepSeek-R1 used RL to directly incentivize long-chain reasoning in large models, proving that RL is not just an "alignment tool" but also a "capability amplifier" — the hottest frontier today; see Frontier Progress.
9. The Master Timeline and the Six Tributaries
1. Key milestones by decade
| Decade | Milestone | Significance in one sentence |
|---|---|---|
| 1950s | Bellman's dynamic programming and the MDP formalization | sequential decision-making gets a mathematical language |
| 1959–60s | Samuel's checkers, Widrow's trial-and-error, Minsky poses credit assignment | the problems are formally stated |
| 1970s–80s | ASE/ACE trial-and-error learners; the seeds of TD | algorithm prototypes appear |
| 1988 | Sutton's TD(λ) | a viable method for online prediction |
| 1989/92 | Watkins's Q-learning | the off-policy revolution |
| 1992 | Williams's REINFORCE | the policy gradient line is established |
| 1992–95 | Tesauro's TD-Gammon | neural-net self-play beats humans |
| 1996 | Kaelbling et al.'s RL survey | the field becomes a system |
| 2013 | Mnih's DQN (NIPS version) | deep RL is born |
| 2015 | DQN (Nature version); TRPO | deep value learning and trust-region policy optimization |
| 2016 | AlphaGo defeats Lee Sedol | search × learning cracks Go |
| 2017 | PPO; AlphaZero (Rainbow paper 2017, formally published at AAAI 2018) | stable policy optimization + the culmination of deep RL |
| 2018 | SAC; World Models; QMIX | continuous control and multi-agent advance together |
| 2019 | OpenAI Five; AlphaStar; RND | scaling and new tools for exploration |
| 2020 | MuZero; the RLHF paper | twin milestones: world models + human feedback |
| 2022 | InstructGPT/ChatGPT | RLHF enters the large-model era |
| 2023 | DPO; DreamerV3 | simplified alignment + a general world model |
| 2025 | DeepSeek-R1 and other RL-for-reasoning work | RL moves from alignment to capability training |
2. Where each tributary leads
Every tributary has a complete learning path on this site:
| Tributary | Starting point | Main path on this site |
|---|---|---|
| Dynamic programming | Bellman 1957 | Paper Map → Markov Decision Process |
| Tabular methods | Samuel 1959 → TD/Q-learning | Value-Based Learning |
| Policy gradient | Williams 1992 → PPO | Policy Gradient Methods |
| Deep RL | DQN 2013 → Rainbow → MuZero | Classic Papers, Closely Read |
| Multi-agent | MADDPG 2017 → MAPPO | Multi-Agent Reinforcement Learning |
| RLHF | Stiennon 2020 → DPO 2023 | RLHF and Alignment from Human Feedback |
The right way to study history
Don't memorize history as a story — read it as a "problem archive": every milestone is a response to the previous one's flaw. TD answered "Monte Carlo has to wait for the whole episode"; Q-learning answered "the learning policy must match the behavior policy"; DQN answered "you can't look up high-dimensional states in a table"; PPO answered "TRPO is too complex"; RLHF answered "a pretrained model only continues the text — it doesn't converse." Grasp this "problem → solution" chain and your understanding of the field connects into a web instead of scattering into dots.
Further Reading
- Paper Map — a paper-level refinement of this timeline, with one-line positions for every paper along the six tributaries.
- Classic Papers, Closely Read — close readings of the six papers that changed RL (including "how to answer this in an interview").
- Value-Based Learning — the full technical thread from dynamic programming → MC/TD → DQN, covering the tabular and deep-value tributaries.
- Policy Gradient Methods — the REINFORCE → TRPO → PPO logic chain: the policy gradient tributary.
- RLHF and Alignment from Human Feedback — the full expansion of the most important tributary since 2020.
- AlphaGo and Monte Carlo Tree Search — the complete case study of the most dramatic member of the 2016 "deep RL trinity."
References
- Bellman, R. (1957). Dynamic Programming. Princeton University Press. — the original source of dynamic programming and the Bellman equation.
- Samuel, A. L. (1959). Some Studies in Machine Learning Using the Game of Checkers. IBM Journal of Research and Development, 3(3), 210–229. https://ieeexplore.ieee.org/document/5392560 — the pioneering work of self-play learning.
- Minsky, M. (1961). Steps Toward Artificial Intelligence. Proceedings of the IRE, 49(1), 8–30. — the first formal statement of the credit assignment problem.
- Sutton, R. S. (1988). Learning to Predict by the Methods of Temporal Differences. Machine Learning, 3(1), 9–44. https://link.springer.com/article/10.1007/BF00115009 — the original TD learning paper.
- Watkins, C. J. C. H. & Dayan, P. (1992). Q-learning. Machine Learning, 8(3–4), 279–292. https://link.springer.com/article/10.1007/BF00992698 — Q-learning and its convergence proof.
- Tesauro, G. (1995). Temporal Difference Learning and TD-Gammon. Communications of the ACM, 38(3), 58–68. https://dl.acm.org/doi/10.1145/203330.203343 — the original TD-Gammon paper.
- Mnih, V. et al. (2015). Human-level control through deep reinforcement learning. Nature, 518, 529–533. https://www.nature.com/articles/nature14236 — the Nature version of DQN.
- Silver, D. et al. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529, 484–489. https://www.nature.com/articles/nature16961 — the original AlphaGo paper.
- Schulman, J. et al. (2017). Proximal Policy Optimization Algorithms. https://arxiv.org/abs/1707.06347 — the original PPO paper.
- Ouyang, L. et al. (2022). Training language models to follow instructions with human feedback (InstructGPT). https://arxiv.org/abs/2203.02155 — the milestone that scaled RLHF to large models.
- Rafailov, R. et al. (2023). Direct Preference Optimization. https://arxiv.org/abs/2305.18290 — the original DPO paper.
- Kaelbling, L. P., Littman, M. L. & Moore, A. W. (1996). Reinforcement Learning: A Survey. JAIR, 4, 237–285. https://arxiv.org/abs/cs/9605103 — the classic survey; a panorama of the golden age.