A Brief History of Reinforcement Learning
Seventy years from the Bellman equation to RLHF: dynamic programming lays the foundations, temporal-difference learning and Q-learning, TD-Gammon, the deep RL trinity (DQN/PPO/AlphaGo), and on to LLM alignment. A decade-by-decade table of the key milestones.