Appearance
Close Reading of Classic Papers
One-liner: this page offers a close, paper-by-paper reading of the six milestone papers that changed RL — Bellman 1957 (dynamic programming), Watkins 1989/1992 (Q-learning), Mnih 2015 (DQN), Silver 2016 (AlphaGo), Schulman 2017 (PPO), and Ouyang 2022 (InstructGPT). Each is covered in five acts — background, method, experiments, limitations, and "how to answer in an interview" — making this ideal if you're prepping for interviews or want to genuinely master these algorithms. By the end, you'll be able to talk about these papers at the level of mechanisms in an interview, not just name-drop them.
What these six papers share: each one overturned an assumption. Bellman gave sequential decision-making a mathematical language; Watkins showed you can learn without knowing the model; Mnih showed you can learn from high-dimensional inputs; AlphaGo fused search with learning; PPO made stable updates simple; InstructGPT took RL out of game environments and into language models. Once you understand which assumption each paper removed, you've effectively installed a GPS on the paper map.
1. How to Read the Six: Order and Method
Recommended order: Bellman → Q-learning → DQN → PPO → AlphaGo → InstructGPT. This follows the dependency chain: each paper's machinery builds on concepts from the ones before it (DQN needs Q-learning's off-policy learning; PPO needs the policy gradient).
For each paper — this matches the L3 detail level in the reading paths — start with the abstract and conclusion to pin down "which assumption does this paper overturn," then work through the method section and derive the core equations, and finally check the experiment tables to confirm that overturning the assumption actually paid off. When the math derails you, switch to the three-column note-taking method — keep "problem / method / limitation" in separate columns instead of banging your head on one equation.
2. Bellman, A Markovian Decision Process (1957)
1. Background: Why This Paper Is the Foundation of Everything
Before the 1950s, "decision making" meant single-shot choices (game theory) or deterministic systems (classical control). Bellman's contribution was to elevate sequential decision-making with delayed rewards into a computable mathematical object: the Markov decision process (MDP). Every RL paper since — including 2025's LLM-reasoning RL — uses the notation (S, A, P, R, γ) introduced here.
2. Method: The Bellman Optimality Equation
The precise definitions of the MDP tuple and the Bellman equation live on the MDP concept page; here's the core directly:
text
State value (under policy π):
V^π(s) = Σ_a π(a|s) Σ_s' P(s'|s,a) [ R(s,a,s') + γ V^π(s') ]
Optimality (the heart of value iteration):
V*(s) = max_a Σ_s' P(s'|s,a) [ R(s,a,s') + γ V*(s') ]Three key mechanisms:
- Principle of optimality: any suffix of an optimal policy is itself optimal. That's what makes the problem recursively solvable — and it's where the name "dynamic programming" comes from.
- The Bellman optimality operator T is a contraction mapping*: for any two value functions,
‖T*V₁ - T*V₂‖∞ ≤ γ‖V₁ - V₂‖∞. Repeatedly applying T* therefore converges to a unique fixed point V* — this is the mathematical essence of value-iteration convergence, and the template for Q-learning's convergence proof two sections down. - Value iteration: start from any initial V and repeatedly apply
V ← T*V; the contraction rate is governed by γ.
3. Limitations
You need known transition probabilities P and rewards R, and the state space must be enumerable (tabular methods). The real world satisfies neither condition — which is exactly why the following papers exist.
4. How to Answer in an Interview
Q: Why do people say every RL algorithm is approximating a solution to the Bellman equation? Answer frame: The Bellman optimality equation V* = TV defines what "optimal value" means. With a known environment you solve it directly (DP). With an unknown environment, Monte Carlo replaces the expectation with samples (TD uses one-step samples plus bootstrapping), Q-learning applies max to the Q function, and DQN represents Q with a neural network — all of them are closing in on the same fixed point under the constraint of not knowing the model.
Frequent follow-ups: What does γ do? → It's the discount factor; with γ<1 returns are bounded and the contraction mapping guarantees convergence; the smaller γ, the more myopic the behavior and the faster the convergence. Value iteration vs. policy iteration? → Value iteration approximates V* directly; policy iteration alternates between policy evaluation and policy improvement.
3. Watkins, Q-learning (1989 PhD Thesis / 1992 Journal)
1. Background: Removing the "Known Model" Assumption
Bellman's equations require P and R. Watkins asked: if all you have are samples of the form "tried this action, got this reward," can you still learn the optimal policy? The answer is yes — provided you introduce the Q function and hand it over to stochastic approximation.
2. Method: The Q Function and the Off-Policy Update
Q-learning defines Q(s,a) as the expected return of "take action a in state s, then act optimally forever after." The update rule:
text
Q(s,a) ← Q(s,a) + α [ r + γ max_a' Q(s',a') - Q(s,a) ]
└────────── TD error δ ──────────┘The off-policy revolution lives in this single formula: max_a' uses a target that assumes the next action is optimal, which is completely decoupled from how actions are actually being chosen right now (ε-greedy, random — whatever). In other words:
- The behavior policy (how data is collected) can be arbitrary and exploratory;
- The target policy (what you learn) is the greedy policy;
- Decouple the two → you can learn from logged, offline data (this directly foreshadows offline RL).
Compare on-policy SARSA, which uses Q(s',a') — the action that will actually be taken. Q-learning is the more optimistic, aggressive sibling, and the difference is qualitative in environments like cliff walking; see the SARSA vs. Q-learning comparison in value-based learning.
3. Intuition for the Convergence Proof (One of Watkins & Dayan 1992's Core Contributions)
The paper's convergence result requires: a finite MDP, every state-action pair visited infinitely often, and learning rates satisfying
text
Σ_t α_t = ∞ (so the updates can travel the full distance)
Σ_t α_t² < ∞ (so the noise gets averaged out)The proof intuition has two layers:
- View the update as a "stochastic approximation to the Bellman optimality operator T*": each sample
r + γ max_a' Q(s',a')is an unbiased sample of(T*Q)(s,a); - T* is a γ-contraction (the mechanism from the Bellman section), so the combination of "unbiased samples + contraction" converges to the unique fixed point Q* under stochastic-approximation theory (the Robbins–Monro framework).
An honest note on the proof
The convergence argument in the 1992 paper was a first version of the result, and the field later produced more rigorous treatments of "Q-learning converges under general conditions" (including counterexamples of non-convergence when exploration is insufficient). In an interview, the "contraction mapping + stochastic approximation" intuition is usually enough; if pressed on rigor, you can concede that convergence relies on infinite exploration and learning-rate conditions — which is exactly why in practice you must ensure every state-action pair gets visited.
4. Limitations
- Tabular: the Q table must store every (s,a) — the moment the state space explodes, this breaks down;
- The max operator introduces systematic overestimation (taking a max over noisy estimates inflates the expectation) — the number-one problem that Double DQN later set out to fix;
- It requires sustained exploration (e.g., ε-greedy); without adequate exploration, no convergence.
5. How to Answer in an Interview
Q: Why is Q-learning off-policy? Can it be trained on behavior-cloning data? Answer frame: Off-policy means the target policy (greedy, from max_a') differs from the behavior policy (the ε-greedy one that collects data). The update only needs (s,a,r,s') tuples, so historical and offline data all work as long as state-action coverage is adequate. This is also the essential difference from SARSA.
Frequent follow-ups: Does Q-learning overestimate values? → Yes; the max operator takes the maximum over noisy estimates, systematically inflating them, and Double DQN decouples "selecting the action" from "evaluating it" to ease the problem. Why should ε decay? → Early on you need exploration, later exploitation; decay too fast and you miss visits, too slow and convergence crawls.
4. Mnih et al., Human-level control through deep reinforcement learning (DQN, 2015, Nature)
1. Background: Two Real Problems in Going from Tables to Networks
In 2013 Mnih et al. first published Playing Atari with Deep Reinforcement Learning (arXiv:1312.5602) to demonstrate feasibility; the 2015 Nature version hardened the system into the milestone of deep RL. Moving Q-learning from a table to a neural network immediately runs into two problems that simply didn't exist in the tabular era:
- Sample correlation: with online updates, consecutive frames are highly correlated, so gradient updates oscillate back and forth;
- Bootstrapping instability: the target
r + γ max Qitself depends on the very network being updated — one update changes both the "answer" and the "question," which easily diverges.
2. Method: Two Engineering Tricks
Trick one: experience replay. Store every transition (s,a,r,s') in a replay buffer, and during training sample minibatches uniformly at random instead of consuming transitions in temporal order:
text
Online use (bad): s1→s2→s3→s4... consecutive samples highly correlated, each used once and discarded
Replay (good): randomly draw {s7,a7,r7,s7'} {s3,a3,r3,s3'} {s9,...} breaks correlation + reuseWhat it does: ① breaks temporal correlation so gradients roughly satisfy the i.i.d. assumption; ② reuses each experience many times, improving data efficiency; ③ also dilutes the dominance of recent experience.
Trick two: target network. Keep a frozen target network Q⁻ that computes the target values, and copy weights over from the online network only every C steps (C=10000 in the Nature version):
text
Online update: Q(s,a) ← Q(s,a) + α[ r + γ max_a' Q⁻(s',a') - Q(s,a) ]
└─ target computed with the frozen network ─┘
Every C steps: Q⁻ ← Q (copy weights)Effect: targets stay fixed for stretches of time, so bootstrapping no longer "chases its own tail," and training becomes dramatically more stable.
3. System Details (Nature Version)
| Component | Setup |
|---|---|
| Input | 84×84 grayscale frames, stacked over the most recent 4 frames (to convey motion) |
| Network | 3 conv layers + fully connected, outputting a Q value per action |
| Frame skip | each action repeated over 4 frames (cuts compute) |
| Exploration | ε-greedy, ε annealed linearly from 1.0 to 0.1, then held fixed |
| Optimizer | RMSProp |
| Loss | TD error clipped (Huber/clipped), reducing gradients from outlier samples |
4. Results and Significance
- The Nature version was evaluated on 49 Atari games and beat the average level of human professional players on 29 of them; the 2013 version only validated 7 games, beating the then-SOTA on 6 — the Nature version is the engineering leap "from making learning work to across-the-board superiority."
- Significance: end-to-end (pixels → actions) general game play worked for the first time, and the deep RL era began. DQN is the protagonist of the Atari case page (Atari and Video Games).
5. Limitations
- The action space must be discrete (the output layer has one Q value per action);
- Overestimation persists (eased by Double DQN, but not eliminated);
- Sample efficiency is still poor (millions of frames required);
- Sensitive to hyperparameters and network architecture; reproducibility was a notorious pain point at the time.
6. How to Answer in an Interview
Q: Why does DQN need experience replay and a target network? Answer frame: Replay solves "temporally correlated samples → unstable gradients" plus "poor data utilization"; the target network solves "target values drift along with the network's own updates → bootstrapping divergence." The two tricks turn deep Q-learning from "untrainable" into "trainable."
Frequent follow-ups: How does replay buffer size affect training? → Too large and the samples go stale (old experience turns misleading once the policy has moved on); too small and correlation creeps back; classic values are on the order of 10^5–10^6. What if the target network updates too frequently? → You're back to chasing your own tail, and instability returns.
5. Silver et al., Mastering the Game of Go with Deep Neural Networks and Tree Search (AlphaGo, 2016, Nature)
1. Background: Why Go Was "Unsearchable"
Go has a state space of roughly 10^170 and a branching factor around 250, so brute-force search is mathematically hopeless. For two decades the best Go programs (rule-based, with hand-crafted evaluation) stood an insurmountable distance behind professional players. AlphaGo's strategy was to trade learning for compute: instead of searching every move, learn which moves are worth searching and what the searched positions are worth.
2. Method: Three-Stage Training Plus Search at Inference
Stage one, SL policy network p_σ(a|s): supervised learning. Train a 13-layer CNN on roughly 30 million positions from human expert games (KGS server) to predict "where a strong human would play," reaching about 57% accuracy. What it learns is the prior distribution of human moves.
Stage two, RL policy network p_ρ(a|s): policy-gradient reinforcement learning. Initialize from the SL network, then play against previous versions of itself and run policy gradient with the win rate as the reward. Result: about an 80% win rate against the SL network. What it learns is "winning" rather than "playing like a human" — stronger than the SL network, but with a different distribution (the classic tension between "imitating humans" and "chasing victory"; see policy gradient).
Stage three, value network v_θ(s): a regression network. Generate about 30 million positions via self-play with the RL policy network, and train the network to directly predict "the win rate from this position, from the RL network's self-play perspective," achieving MSE around 0.23 — markedly better than the fast rollouts' roughly 0.48.
text
Three-stage training pipeline:
human game records ──SL──> p_σ (learns "how humans play") ──policy gradient──> p_ρ (learns "how to win")
│ self-play positions
▼
v_θ (learns "position value")
At inference: in MCTS, p_ρ provides prior probabilities and v_θ evaluates leaf nodes,
the two combine into the search's guiding signal — "search × learning" convergeAt inference: MCTS. The four steps of Monte Carlo tree search (selection/expansion/simulation/backpropagation) and the UCB variant are detailed in the AlphaGo and MCTS case study. AlphaGo's distinctive touch: every node in the tree gets a prior P(a|s) from the policy network, and the selection score scales with the prior (the moves worth exploring have already been shortlisted by the network); leaf nodes are evaluated by a blend of the value network and fast rollouts. The match version ran distributed across about 1202 CPUs and 176 GPUs; the single-machine version used about 48 CPUs and 8 GPUs.
3. Results and Significance
- In October 2015 it beat European champion Fan Hui 5:0 — the first computer program to defeat a professional player;
- In March 2016 it beat world-leading player Lee Sedol 4:1 (Lee won game 4 with the famous "divine move," an all-time moment in Go history).
- Significance: not just "AI beat humans at Go," but fusing three schools — supervised learning, reinforcement learning, and tree search — into a single system. None of the three alone could have beaten a top player; together they could. This is widely regarded as the landmark event of deep RL and AI.
4. Limitations
- Training cost is extreme (thousands of GPU/TPU-equivalents of compute), not generalizable to everyday tasks;
- The SL stage depends on human game records (AlphaZero later removed it via pure self-play);
- The system-engineering complexity of dual networks plus distributed rollouts is immense.
5. How to Answer in an Interview
Q: What problem does each of AlphaGo's three training stages solve? Answer frame: The SL network solves "initialization and priors" (where strong humans play, giving search a sensible starting point); the RL network solves "going from human-like to winning" (self-play reinforcement pulls it well beyond the SL network); the value network solves "position evaluation" (regressing wins and losses on self-play data replaces expensive rollouts). At inference, MCTS fuses "policy priors × value evaluation" into the search signal — the products of learning become the inputs to search.
Frequent follow-ups: Why not train the value network on human game records? → The win/loss distribution of human records diverges from the RL network's actual game distribution; training on RL self-play data keeps the evaluator aligned with real use. What did AlphaZero change relative to AlphaGo? → Dropped the SL stage and hand-crafted features, learning purely from self-play starting from scratch; see the paper map and model-based RL.
6. Schulman et al., Proximal Policy Optimization Algorithms (PPO, 2017)
1. Background: TRPO Is Stable but Complicated — and Painful in Practice
TRPO (2015, arXiv:1502.05477) proved that limiting the update size (a trust region) stabilizes policy gradients, but it requires solving a constrained quadratic program with a KL constraint: Fisher information matrix, conjugate gradients, line search — complicated to implement, painful to tune, and hard to parallelize at scale. PPO's question: can first-order optimization approximate the trust region and stay both stable and simple?
2. Method: The Clipped Objective
Write the probability ratio of new to old policy as r_t(θ) = π_θ(a_t|s_t) / π_θold(a_t|s_t) (θold being the parameters in force when the data was collected). PPO's objective:
text
L^CLIP(θ) = E_t[ min( r_t(θ) Â_t, clip(r_t(θ), 1-ε, 1+ε) Â_t ) ], ε=0.2
Intuition, piece by piece:
· Â_t is the advantage estimate (computed with GAE — "how much better than average was this action")
· When the advantage is positive (a good action): once r_t exceeds 1+ε the objective is clipped → no extra credit for "one update stepping too far"
· When the advantage is negative (a bad action): once r_t drops below 1-ε it's likewise clipped → a single bad sample can't crush the probability
· The min() guarantees: if the unclipped objective is lower, the unclipped one is used → the objective can never be artificially inflatedThe key point: clip constrains the probability ratio between old and new policies (a relative change), not parameter distance or action distance. That gives PPO a "trust region" effect with nothing but first-order gradients.
The full algorithm (per iteration):
text
1. Collect a batch of trajectories with the current policy π_θold
2. Compute the advantage Â_t at every timestep with GAE (see [policy gradient](/concepts/policy-gradient))
3. Run K epochs of minibatch gradient ascent on the same batch to maximize L^CLIP
4. Update the critic with TD error; θold ← θ3. Why It Replaced TRPO (the Answer from the Paper and Later Practice)
| Dimension | TRPO | PPO |
|---|---|---|
| Optimization | second-order (Fisher matrix + conjugate gradient) | first-order (plain gradient ascent) |
| Constraint form | hard KL constraint | clip as a soft constraint (approximate) |
| Implementation complexity | high (line search, matrix ops) | low (a few lines of code) |
| Parallelization / scale | awkward | naturally suited (data reused for multiple epochs) |
| Stability | strong (strict constraint) | strong (clip is close enough) |
The paper validated PPO on Atari and MuJoCo: sample efficiency and stability match or exceed TRPO, A2C, and other mainstream methods of the day, while the implementation is an order of magnitude simpler. It also became the workhorse optimizer for RLHF fine-tuning of language models at OpenAI and elsewhere — the PPO fine-tuning stage in the LLM alignment case study is exactly this.
4. Limitations
- Still on-policy (each batch is used once and discarded), so sample efficiency lags off-policy methods;
- The clip ε needs tuning: too large and the approximation breaks down; too small and steps get timid;
- Sensitive to reward scale (in RLHF this is handled with a KL penalty; see RLHF).
5. How to Answer in an Interview
Q: What exactly does PPO's clip constrain, and why can it replace TRPO's hard constraint? Answer frame: It constrains the probability ratio r_t(θ) between old and new policies — i.e., the size of each update step. When a single sample pushes the ratio past [1-ε, 1+ε], the objective gets clipped, preventing one oversized update. Because a small deviation in the probability ratio approximately corresponds to a small KL distance inside the trust region, clip is a first-order implementable approximation of the trust region — TRPO-level stability at an order-of-magnitude lower complexity.
Frequent follow-ups: Why the min-plus-clip combination instead of a plain clip? → To prevent the clipped objective from being artificially inflated: when the unclipped objective is lower, optimization follows the true value, so the policy isn't misled where it stands to gain. Why ε = 0.2? → An empirical value from the paper; anything in 0.1–0.3 is stable on Atari/MuJoCo. Is PPO on-policy or off-policy? → On-policy (behavior and target are the same policy), though the data can be reused for several epochs of minibatch updates.
7. Ouyang et al., Training Language Models to Follow Instructions with Human Feedback (InstructGPT, 2022, arXiv:2203.02155)
1. Background: Pretrained Models "Can Continue Text but Can't Do the Work"
GPT-3 can continue any text, but hand it an instruction and it may not comply — it might answer the wrong question, fabricate facts, or produce harmful content. The reason: pretraining optimizes "predict the next token," not "satisfy the user's intent." Behavior cloning (SFT) can teach "format" but not "what counts as a good answer" — because "good" is subjective and can't be exhaustively labeled. InstructGPT's answer: turn "human preference" into a reward signal and optimize it with RL.
2. Method: The Three-Stage Pipeline
Stage one, SFT (supervised fine-tuning): collect about 13k prompts (real prompts from OpenAI API users), have labelers write the desired responses, and fine-tune GPT-3. The result is the SFT model. Limitation: what it learns is "answers like the labelers'," not "answers users prefer," and the sample count is small.
Stage two, the reward model (RM): for each prompt, have the SFT model generate 4–9 responses, and have labelers rank them (rather than score — ranking is more consistent). Fit the rankings with a Bradley–Terry model as a logistic loss:
text
loss = -E[ log σ( r_θ(x, y_w) - r_θ(x, y_l) ) ]
# y_w is the response ranked higher, y_l the one ranked lower
# r_θ assigns a scalar score to "how good this response is"This RM turns "human intuition" into a differentiable reward function. The mechanics of preference ranking are covered in the Bradley–Terry section of the RLHF concept page.
Stage three, PPO fine-tuning: use the RM's output as the reward and run PPO on the SFT model. The critical engineering detail — reference model and KL penalty. The RLHF objective:
text
maximize E[ r_φ(x, y) ] - β · KL( π_θ(y|x) ‖ π_SFT(y|x) )
# please the reward model, but don't drift too far from the SFT model (β on the order of 0.02)The KL penalty keeps the policy from "exploiting loopholes in the reward model" (overoptimization; see reward engineering). The paper also mixes in a small amount of pretraining gradient (ppo-ptx) so the model doesn't lose its general language ability.
3. Results (Human Evaluation)
- The 1.3B InstructGPT beats the 175B GPT-3 in human evaluation — not "smaller is better," but "the gains from alignment outweigh the gap in scale";
- Outputs are more "helpful" (better instruction-following), more honest (noticeably fewer hallucinations on TruthfulQA), and improved on harmfulness;
- Alignment tax: average performance on standard academic benchmarks (SQuAD, HellaSwag, WMT, etc.) drops slightly (about 0.4%) — "pleasing human preferences" and "optimizing benchmark scores" don't perfectly align.
A commonly misunderstood point
The reward in InstructGPT's "RLHF" is not some game score — it's the reward model's rating, and the reward model is itself learned and biased. So RLHF carries, from first principles, the risk of "reward model is wrong → the policy follows the wrong reward." This is the frontier problem discussed in reward overoptimization.
4. Limitations
- Human labeling is expensive; the scale and quality of preference data set the ceiling;
- The reward model is a proxy objective, and RLHF training is sensitive to the preference distribution;
- The paper's own evaluators are drawn from the labeler pool, introducing subjectivity and potential bias;
- The three-stage pipeline is complex engineering (five models on stage at once: SFT model + RM + reference model + policy + critic); see the hands-on breakdown in the LLM alignment case study.
5. How to Answer in an Interview
Q: Why doesn't InstructGPT use pure SFT, or do supervised learning directly on the reward? Answer frame: SFT teaches "format," not "quality"; the reward model turns human preference into an optimizable objective; and PPO performs policy optimization against this non-differentiable reward. The RLHF three-stage pipeline is the complete chain of "human judgment → differentiable reward → policy optimization." The KL penalty is the key guardrail: it keeps the policy from drifting too far from the SFT model, and it's the starting point for the later research on reward overoptimization.
Frequent follow-ups: Why train the RM with rankings instead of scores? → Humans are more consistent about relative goodness, and rankings agree better across labelers. What happens if the KL penalty β is too large? → The policy barely moves and alignment fails; too small → the reward model gets overoptimized and outputs collapse into mode degeneration. Difference between DPO and InstructGPT? → DPO uses the closed-form solution of the preference objective to skip the RM and PPO entirely; see RLHF.
8. The Six Papers as One Line: A 5-Minute Review Before the Interview
text
Bellman 1957 defines "optimal" V* = max [R + γV*] (contraction mapping)
↓
Watkins 1992 learn with an unknown model Q update = stochastic approximation + contraction; off-policy
↓
Mnih 2015 learn from high-dimensional input replay breaks correlation + target network stabilizes bootstrapping
↓
Schulman 2017 stable updates made simple clip the probability ratio ≈ first-order trust region
↓
Silver 2016 learning meets search three-stage training + MCTS; trading learning for compute
↓
Ouyang 2022 RL leaves environments, enters three stages (SFT → RM → PPO+KL)
languageWalk through this chain once before the interview — for each link, spell out "what gap the previous link left and how this one filled it" — and you'll be ready for any "why" question.
Further Reading
- Paper Map — places all six papers in a global coordinate system of six tributaries and a seventy-year timeline.
- Value-Based Learning — mechanism-level expansion of Bellman/Q-learning/DQN (DP, MC/TD, SARSA, Rainbow).
- Policy Gradient — mechanism-level expansion from REINFORCE to PPO (GAE, baseline, clip).
- RLHF and Alignment with Human Feedback — mechanism-level expansion of InstructGPT's three stages (Bradley–Terry, KL, DPO).
- AlphaGo and Monte Carlo Tree Search — the full AlphaGo case study: MCTS's four steps, three-stage training, and the AlphaZero/MuZero evolution.
- LLM Alignment: RLHF in Practice — the engineering view of InstructGPT: five models on one stage, the alignment tax, and open-source reproductions.
References
- Bellman, R. (1957). A Markovian Decision Process. Indiana University Mathematics Journal 6(4):679–684. https://doi.org/10.1512/iumj.1957.6.56038
- Watkins, C. J. C. H. (1989). Learning from Delayed Rewards (PhD thesis). University of Cambridge. https://www.cs.rhul.ac.uk/~chrisw/new_thesis.pdf — the original dissertation, publicly available.
- Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine Learning 8(3):279–292. https://link.springer.com/article/10.1007/BF00992698
- Mnih, V., et al. (2013). Playing Atari with Deep Reinforcement Learning. arXiv:1312.5602. https://arxiv.org/abs/1312.5602
- Mnih, V., et al. (2015). Human-level control through deep reinforcement learning. Nature 518:529–533. https://www.nature.com/articles/nature14236
- Silver, D., et al. (2016). Mastering the Game of Go with Deep Neural Networks and Tree Search. Nature 529:484–489. https://www.nature.com/articles/nature16961
- Schulman, J., et al. (2015). Trust Region Policy Optimization. arXiv:1502.05477. https://arxiv.org/abs/1502.05477
- Schulman, J., et al. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347. https://arxiv.org/abs/1707.06347
- Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback (InstructGPT). arXiv:2203.02155. https://arxiv.org/abs/2203.02155
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. https://incompleteideas.net/book/RLbook2020.pdf