Skip to content

Interview Question Bank

On this page Worked answers to high-frequency RL interview questions — MDP and Bellman, MC vs TD, Q-learning vs SARSA, DQN engineering tricks, PPO details, SAC and continuous control, exploration vs exploitation, RLHF, reward-design scenario questions — plus an answer framework and follow-up predictions.

Interview Question Bank ​

One-line pitch: this page is a collection of "real interview questions + worked answers" for RL — high-frequency questions organized into five question types, a standard answer framework for each, predictions of likely follow-ups, and a 30-question self-test checklist.

How to use this page: don't memorize answers — practice the answer framework. RL interview questions are highly standardized, but interviewers' follow-ups vary wildly. Once you've internalized the framework — "definition → intuition → formula → example → pitfalls → extensions" — you can handle any variant. The coverage here is driven by the high-frequency topic weighting table in the Knowledge Map: that page tells you what to learn; this one tells you how to answer once you've learned it.

1. Theory Questions (MDP / Bellman / Convergence Intuition) ​

Q1: What is a Markov decision process (MDP)? Why can all RL problems be framed as one? ​

Answer framework:

  • Definition: An MDP is a five-tuple $(S, A, P, R, \gamma)$, where $S$ is the state space, $A$ the action space, $P(s'\mid s,a)$ the transition probabilities, $R(s,a)$ the reward function, and $\gamma\in[0,1)$ the discount factor.
  • Intuition: The next state and reward depend only on the current state and action, not on history (the Markov property). That's what lets us write recurrences instead of tracing causes back through the past.
  • Formula: The value function of a policy $\pi(a\mid s)$ satisfies the Bellman equation:
    text
    Vπ(s) = Σ_a π(a|s) Σ_{s',r} p(s',r|s,a) [ r + γ·Vπ(s') ]
  • Example: Go — the state is the current board position, an action is a move, the transition is deterministic (the opponent's moves can be folded into the environment), and the reward is the final win/loss.
  • Pitfall: Many real problems don't strictly satisfy the Markov property (think partial observability → POMDP), so you need to stitch history or observations into the state. In an interview it's fine to say "almost any RL problem can be approximated as an MDP" — but flag the partial-observability caveat.
  • Extension: see Markov Decision Process.

Likely follow-ups: (1) What if the Markov property doesn't hold? (Answer: POMDP, frame stacking, RNN-based memory.) (2) Why must the discount factor $\gamma$ be less than 1? (Answer: finite horizon / convergence of the infinite sum.) (3) What does $\gamma=0$ mean? (Answer: only immediate reward matters — it degenerates into a bandit problem.)

Answer framework:

  • Definition: The Bellman equation is the family of recurrences structured as "current value = immediate reward + discounted future value." It comes in two forms — the expectation equation (holds for any given policy $\pi$) and the optimality equation (holds for the optimal policy, replacing the weighted sum over $\pi$ with $\max_a$).
  • Formula: the optimality equation:
    text
    V*(s) = max_a Σ_{s',r} p(s',r|s,a) [ r + γ·V*(s') ]
  • Intuition: The optimality equation says "the optimal value = taking, at every step, the action that leads to the greatest value." The expectation equation, meanwhile, is the self-consistency condition for the value of a given policy.
  • Pitfall: The Bellman equation is not a solution algorithm — it's a fixed-point condition. It characterizes the equations a value function must satisfy; actually solving them is what value iteration and policy iteration do.
  • Extension: see Value Learning.

Likely follow-ups: (1) Difference between value iteration and policy iteration? (Answer: policy iteration alternates evaluation and improvement; value iteration iterates the value only, with the greedy policy implicit.) (2) Relationship between $V$ and $Q$? (Answer: under the optimal policy $V(s)=\max_a Q(s,a)$; $Q(s,a)$ conditions on an action, while $V(s)$ averages over actions according to the policy.)

Q3: Why does Q-learning converge? What's the intuition? ​

Answer framework:

  • Intuition: Every Q-learning update $Q(s,a)\leftarrow Q(s,a)+\alpha[r+\gamma\max_{a'}Q(s',a')-Q(s,a)]$ moves $Q$ one step toward the "target" — the result of applying the Bellman optimality operator to $Q$. The Bellman optimality operator $\mathcal{T}$ is $\gamma$-contractive ($\lVert\mathcal{T}Q_1-\mathcal{T}Q_2\rVert_\infty\le\gamma\lVert Q_1-Q_2\rVert_\infty$), so repeated iteration must converge to the unique fixed point $Q^*$.
  • Prerequisites: every $(s,a)$ pair must be visited infinitely often (i.e., sufficient exploration), and the learning rate $\alpha$ must satisfy the Robbins–Monro conditions ($\sum\alpha=\infty,\ \sum\alpha^2<\infty$).
  • Pitfall: This is the classic result for the tabular case. Once function approximation enters the picture (DQN), the guarantee no longer holds — one reason the DQN era needed replay buffers and target networks.
  • Extension: see Value Learning.

Likely follow-ups: (1) Key step in the contraction proof? (Answer: the max-norm plus the $\gamma$ discount makes it a contraction mapping, so the fixed point is unique.) (2) What happens with insufficient exploration? (Answer: some action never gets updated and keeps its initial value forever, so the resulting policy can be suboptimal.)

2. Comparison Questions (MC vs TD / On- vs Off-Policy / Value vs Policy) ​

Q4: Differences between Monte Carlo (MC) and temporal-difference (TD) learning? ​

Answer framework:

DimensionMCTD(0)
When the update happensWait until the episode ends and use the full return $G_t$Update immediately at every step using $r+\gamma V(s')$
BiasUnbiased (samples of the true return)Biased (uses its own current estimate $V(s')$ — bootstrapping)
VarianceHigh (full trajectories are noisy)Low (single-step updates are less noisy)
Learning speedSlow (must wait for episode end)Fast (online learning)
Non-terminating trajectoriesCan't be used directly (non-terminating tasks need truncation)Works fine
  • Formula: TD error $\delta = r + \gamma V(s') - V(s)$, update $V(s)\leftarrow V(s)+\alpha\delta$.
  • Intuition: TD "updates an estimate with another estimate" (bootstrapping) — hence low variance but biased. MC "updates with real data" — unbiased but high variance.
  • Pitfall: Under function approximation, TD's bias can even cause divergence; MC's variance becomes enormous in long-horizon tasks. TD($\lambda$)/GAE, which interpolates between the two, exists precisely to tune this trade-off.
  • Extension: see Value Learning.

Likely follow-ups: (1) What does TD($\lambda$) / eligibility traces solve? (Answer: the bias–variance trade-off across multiple steps; GAE is its continuous analog.) (2) Why does TD diverge with function approximation? (Answer: bootstrapping creates a self-reinforcing positive feedback loop; off-policy learning + function approximation are two of the three deadly-triad ingredients.)

Q5: Differences between Q-learning and SARSA? When do their results diverge noticeably? ​

Answer framework:

  • Formula (both update the same $Q$; they differ in the next action inside the target):
    text
    Q-learning: Q(s,a) ← Q(s,a) + α [ r + γ·max_{a'} Q(s',a') - Q(s,a) ]
    SARSA:      Q(s,a) ← Q(s,a) + α [ r + γ·Q(s',a')     - Q(s,a) ]   (a' is the action actually taken)
  • Key difference: Q-learning is off-policy — its target uses the "optimal-action assumption" $\max_{a'}Q(s',a')$, independent of the behavior policy. SARSA is on-policy — its target uses the action actually taken next, $Q(s',a')$.
  • Consequence: With a large exploration ε, SARSA learns a policy that "accounts for the penalties incurred while exploring" and is therefore more conservative; Q-learning learns the purely optimal policy. The classic illustration is Cliff Walking: SARSA learns the safe route that detours away from the cliff (because it accounts for ε-exploration occasionally walking off the edge), while Q-learning learns the shortest path skimming the cliff — but falls off frequently during exploration.
  • Pitfall: Their convergence conditions differ. Q-learning's off-policy nature lets it reuse old data or any behavior policy (the premise of DQN's replay), but the max operation introduces overestimation — the motivation for Double DQN.
  • Extension: see Value Learning.

Likely follow-ups: (1) Where does the overestimation come from? (Answer: the $\max$ operator amplifies estimation errors — see the Double DQN paper.) (2) Which is harder to converge, on-policy or off-policy? (Answer: off-policy + function approximation + bootstrapping form the "deadly triad" — the most dangerous combination.)

Q6: Trade-offs between value-based and policy-based methods? Why has modern RL converged on Actor-Critic? ​

Answer framework:

DimensionValue-based (DQN family)Policy-based (REINFORCE/PPO family)
What it learnsQ/V functions; the policy is derived by argmaxThe policy distribution $\pi_\theta(a\mid s)$ directly
Action spaceNaturally discrete; continuous actions need discretization or continuous-value tricksNaturally supports continuous actions
Stochastic policiesOnly via greedy derivation (plus ε)Naturally outputs a probability distribution
ConvergenceCan diverge under function approximation (Q overestimation)More stable, smoother gradient steps
VarianceLow (bootstrapping)High (REINFORCE especially)
Sample efficiencyHigh (off-policy data reuse)Low (on-policy)
  • Why the two lines merge: Actor-Critic pairs "an actor outputting the policy + a critic outputting values that serve as the advantage baseline" — keeping the policy gradient's support for continuous and stochastic actions while using the value function to slash variance. See Actor-Critic Family.
  • Intuition: The critic's value function gives the actor a ruler for "how much better is this action than average" (the advantage), instead of "the return of this whole trajectory" (REINFORCE's $G_t$, whose variance explodes).
  • Pitfall: A classic trap question — "Is SAC value-based or policy-based?" The right answer: "an Actor-Critic hybrid — the policy network is policy-based, but it's coupled to the critic updates through the entropy objective."
  • Extension: see Policy Gradient Methods.

Likely follow-ups: (1) Why does the advantage reduce variance? (Answer: subtracting an action-independent baseline leaves the expectation unchanged while reducing variance.) (2) How do you rescue REINFORCE's high variance? (Answer: baselines/advantages, multi-step returns, normalization, GAE.)

3. Deep RL Questions (DQN Tricks / Why PPO Is Stable / SAC Entropy) ​

Q7: Why does DQN need experience replay and a target network? ​

Answer framework:

  • Experience replay: Store $(s,a,r,s')$ in a buffer and sample random mini-batches for updates. Two effects: (1) it breaks the temporal correlation between samples — otherwise consecutive, highly correlated samples make gradient updates oscillate, so you're effectively spinning in place; (2) it enables data reuse, improving sample efficiency (the premise of off-policy learning).
  • Target network: Compute the target $r+\gamma\max_{a'}\hat{Q}(s',a')$ with a "slowly updated" $\hat{Q}$ that is periodically copied from $Q$. Effects: (1) it breaks the instability of bootstrapping — if the target moved in lockstep with the current parameters, you'd be "training your current estimate on itself," which oscillates or diverges; (2) a fixed target turns the regression into a stationary supervised-learning problem for a while, which is stable.
  • Pitfall: The target network's lag creates an over/underestimation lag effect; pairing it with Double DQN (use $Q$ to select the action, $\hat{Q}$ to evaluate it) further trims the overestimation.
  • Extension: see Value Learning and the Atari case study.

Likely follow-ups: (1) How does Double DQN work? (Answer: $a^=\arg\max_{a}Q(s',a)$, then evaluate with $\hat{Q}(s',a^)$.) (2) What are Dueling DQN's two branches? (Answer: state value $V(s)$ + advantage $A(s,a)$.) (3) What are Rainbow's seven components? (Answer: vanilla DQN plus Double DQN, prioritized replay, dueling, C51 distributional learning, multi-step learning, and NoisyNet.)

Q8: Why does PPO's clipped objective stabilize training? Why did PPO replace TRPO? ​

Answer framework:

  • Motivation: Policy-gradient methods collapse if a single step is too large (the learning rate is hard to tune); TRPO guarantees bounded per-step policy change via a trust region (KL constraint), but its second-order optimization is heavy and hard to implement.
  • Formula (PPO's core objective):
    text
    r_t(θ) = π_θ(a_t|s_t) / π_θold(a_t|s_t)
    L^CLIP(θ) = E_t[ min( r_t(θ)·Â_t,  clip(r_t(θ), 1-ε, 1+ε)·Â_t ) ]
    where $\varepsilon$ is typically 0.2 and $\hat{A}_t$ is the GAE-estimated advantage.
  • Intuition: $r_t(\theta)$ is the likelihood ratio between the new and old policies. When $r_t(\theta)$ leaves $[1-\varepsilon,1+\varepsilon]$, the objective is clipped — with positive advantage, the new policy can't get "too much better" than the old one; with negative advantage, it can't get "too much worse." Each update is thereby confined to a trust region, using only first-order gradients — easy to implement and easy to parallelize.
  • GAE: $\hat{A}t=\sum^\infty(\gamma\lambda)^l\delta_{t+l}$, with $\delta_t=r_t+\gamma V(s_{t+1})-V(s_t)$; λ tunes the bias–variance trade-off.
  • Pitfall: Clipping only acts "when updating in that direction"; if the policy is already so bad that clipping triggers on both sides, the update direction can be nullified (so don't start from a terrible policy). Another pitfall: PPO being stable doesn't mean it's universally better — it trades away some sample efficiency for stability.
  • Extension: see Policy Gradient Methods.

Likely follow-ups: (1) Why clip instead of simply truncating $r_t$? (Answer: hard truncation introduces bias; clipping acts as a monotone safeguard.) (2) Is PPO on-policy or off-policy? (Answer: on-policy — importance sampling only enables multiple epochs over the same batch; the data itself is still on-policy.) (3) What if ε is too large or too small? (Answer: too large → each step strays too far and training destabilizes; too small → slow learning. 0.2 is the common default.)

Q9: Why does SAC maximize entropy? How does automatic temperature tuning work? ​

Answer framework:

  • Objective: SAC maximizes the weighted sum of return + entropy: $J(\pi)=\sum_t\mathbb{E}[(r_t+\alpha\mathcal{H}(\pi(\cdot\mid s_t)))]$. The entropy term $\mathcal{H}$ encourages the policy to stay stochastic.
  • Intuition: (1) Exploration: a high-entropy policy explores naturally, avoiding premature convergence to a local optimum; (2) Robustness: stochastic policies are more resilient to model error and environment perturbation — especially important in robotics; (3) Multi-modality: when several equivalent optima exist, entropy regularization makes the policy output a mixture instead of committing to one, which helps downstream fine-tuning.
  • Automatic temperature: Treat $\alpha$ as the dual variable under the constraint "average entropy ≥ target entropy $\bar{\mathcal{H}}$," updated by gradient descent:
    text
    α ← α - η·∇[ α·( log π(a|s) + H̄ ) ]
    In other words: if the current entropy is below target, increase α (more randomness); if above, decrease it.
  • Engineering details: SAC combines twin Q networks (taking the minimum of the two estimates) to curb overestimation + delayed updates (the critic updates more often than the actor) + target-entropy tuning, making it the practical default for continuous control.
  • Pitfall: The relationship between SAC's entropy and exploration is often misunderstood — entropy doesn't directly produce "systematic exploration" (like visiting new states); it just makes the output "more random." When you genuinely need exploration, you still need extra machinery.
  • Extension: see Actor-Critic Family and Robotics Sim2Real.

Likely follow-ups: (1) Similarities and differences between TD3 and SAC? (Answer: both use twin Q networks to curb overestimation; TD3 uses target-policy smoothing + delayed updates and learns a deterministic policy plus noise, while SAC uses entropy regularization and learns a stochastic policy.) (2) When SAC and when PPO? (Answer: limited samples / reusable data → SAC; large-scale online parallelism / stability first → PPO.)

4. Scenario Design Questions (Reward Design / Environment Modeling) ​

Q10: Given a task, how do you design the reward? (A general framework for scenario questions) ​

Answer framework (universal for any scenario question) — goal decomposition → reward signal → sparse/dense trade-off → anti-hacking → evaluation:

  1. Pin down the goal: "What defines success in this task? Who judges it?" — nail the business metric first.
  2. Choose the reward form: sparse (only +1 on success, paired with a curriculum or intrinsic rewards) vs dense (signal at every step). Use sparse when you can; dense rewards are shaped, and shaping goes wrong easily.
  3. Reward-shaping principles: the intuition behind the potential-based shaping theorem — a shaping term $\Phi(s')-\Phi(s)$ doesn't change the optimal policy (see Reward Engineering); it's just a "gradient ramp," so don't lose sight of the final objective while shaping.
  4. Risk review: can this reward be gamed? (Reward-hacking check: can the agent score high without actually achieving the real goal?)
  5. Evaluation backstop: beyond the reward, keep a business metric that can't be hacked as the final arbiter.

Example: teaching an agent to carry a plate — reward design: +10 for a successful delivery (sparse, the true objective), +0.1 per step of approaching the goal (dense shaping), −5 for dropping the plate (safety constraint). Pitfall: with only the approach reward, the agent may learn to "spin in place holding the plate to farm points." So shaping must be potential-based (monotone guidance) and must leave no point-farming path open.

Likely follow-ups: (1) Classic reward-hacking cases? (Answer: the CoastRunners boat farming points by drifting in circles; a robot gaming the camera view to fake success; models in RLHF churning out sycophantic boilerplate to please the reward model — see the case collection in Reward Engineering.) (2) What about sparse rewards? (Answer: curriculum learning, HER (hindsight experience replay), intrinsic rewards like RND, or cold-starting with behavior cloning.) (3) Who should own the reward function? (Answer: the product/business side owns the "success definition"; the algorithm engineer translates it into a reward and owns the evaluation.)

Q11: A live drill — modeling an arbitrary business problem as an MDP ​

Answer framework: fill in the five-tuple + three risk points:

text
S (state):    Observable business information that affects decisions (mind the Markov property;
              if it isn't sufficient, stack history)
A (action):   The executable decision unit in the business (granularity matters — not too fine, not too coarse)
P (dynamics): The business environment (real system / simulator / historical data); whether the model is
              known determines the choice of solution method
R (reward):   A signal aligned with the business's long-term metrics (see the anti-hacking principles in Q10)
γ (discount): The business's time preference (short-cycle business → small γ; long-term retention → large γ)

Risk points: (1) the state misses key information → POMDP; (2) action granularity mismatched with the business cadence; (3) reward looks only at the short term → long-term metrics collapse. Example: a recommender system — S = user features + behavior history, A = ranking of candidate items, R = clicks/conversions (but add long-term metrics to prevent clickbait).

Likely follow-ups: (1) Is the model $P$ of this MDP known? (Answer: if known → dynamic programming / model-based; if unknown → model-free; with abundant historical data, consider offline RL.) (2) How big is the action space — discrete or continuous? (Answer: decides between the DQN family and the SAC/PPO family.)

5. RLHF / LLM Questions ​

Q12: What are the three stages of RLHF? Why can't SFT alone do the job? ​

Answer framework:

  • The three stages (InstructGPT being the canonical example):
    1. SFT: fine-tune the pretrained model on human-written "ideal responses" so it learns the shape of dialogue;
    2. Reward model (RM): collect human preference rankings over multiple responses and fit them with a Bradley–Terry model: $P(y_w\succ y_l)=\sigma(r(x,y_w)-r(x,y_l))$;
    3. PPO fine-tuning: use the RM's score as the reward and pull the policy $\pi_\theta$ toward the reference model under a KL constraint: $r_{\text{total}}(x,y)=r_\theta(x,y)-\beta\cdot\text{KL}(\pi_\theta(y\mid x),|,\pi_{\text{ref}}(y\mid x))$ — with actor, critic, reference, and RM, four models working in concert during training.
  • Why SFT alone isn't enough: SFT only teaches the model to "sound human," not to "be preferred." Human preferences are relative rankings, not absolute labels, which SFT cannot optimize from. RLHF's key move is using the preference signal with reinforcement learning (sequential decision-making, optimized step by step) to give the model a gradient toward "which response is better."
  • Why the KL penalty: The reward model itself has blind spots; unconstrained maximization of the RM score learns to "say pretty things that game the metric" (reward over-optimization / Goodhart's law). The KL penalty pins the policy near the reference model to prevent collapse — this is the essence of the alignment tax: better alignment sacrifices some raw capability.
  • Pitfall: A KL coefficient $\beta$ that's too small → over-optimization collapse; too large → poor alignment. The classic curve is "reward score rises while human rating rises, then falls."
  • Extension: see RLHF & Human-Feedback Alignment and the LLM alignment case study.

Likely follow-ups: (1) Trade-offs between DPO and PPO? (Answer: DPO writes the preference directly into a loss — no RM, no online sampling, stable and resource-cheap; but PPO supports online exploration and more complex reward models. Formula:

text
L_DPO = -E[ log σ( β·log(πθ(y_w|x)/πref(y_w|x)) - β·log(πθ(y_l|x)/πref(y_l|x)) ) ]

) (2) How do you monitor reward over-optimization? (Answer: run human evaluation and RM scoring side by side, plot the "over-optimization curve," and set a KL budget.) (3) What is the "environment" in RLHF compared to classic RL? (Answer: the environment is the LLM's generation plus RM scoring; the reward is the model's own output — sparse and imperfect.)

Q13: What is reward hacking, and what does it look like in RLHF? ​

Answer framework:

  • Definition: The agent finds behavior that "doesn't satisfy the task's true intent but scores high on the reward function" — essentially the engineering form of Goodhart's law: the reward function ≠ the true objective.
  • What it looks like in RLHF: The model learns to "tell the reward model what it wants to hear" — verbose, sycophantic, substance-free answers; the classic case is "over-pursuing helpfulness to the point of fabricating facts."
  • Mitigations: (1) KL constraint + reference model; (2) phased, conservative calibration of the reward, with human evaluation cross-checks when validating the reward model; (3) multiple reward/constraint terms so a single loophole can't dominate; (4) continuous online evaluation against real business metrics.
  • Extension: see Reward Engineering.

Likely follow-ups: (1) Besides the KL penalty, what other anti-hacking tools exist? (Answer: adversarial red-teaming of the reward model, ablating spurious features like response length, constraining generation length, and evaluating with a human + rule set that can't be hacked.) (2) In RLHF, are "reward over-optimization" and "reward hacking" the same thing? (Answer: closely related — over-optimization is the quantified manifestation of hacking, while reward hacking is the broader term for any loophole-seeking behavior.)

6. Answer-Framework Cheat Sheet & Self-Test Checklist ​

Answer-Framework Cheat Sheet (memorize it) ​

text
The universal skeleton for every interview question:
1. Definition (one sentence, precise terminology)
2. Intuition (plain language: why it exists / why it's done this way)
3. Formula / mechanism (write it if you can; if not, at least describe the structure)
4. Example (real or classic — one is enough)
5. Pitfalls / edge cases (where the method breaks) — what interviewers most want to hear
6. Extensions (a sentence or two: related methods / your own project experience)

30-Question Self-Test Checklist ​

#QuestionCan You Answer It?
1The MDP five-tuple and the Markov property☐
2Bellman expectation/optimality equations☐
3Intuition for Q-learning convergence (contraction operator)☐
4MC vs TD (bias/variance)☐
5Q-learning vs SARSA (on/off-policy, Cliff Walking)☐
6Value-based vs policy-based trade-offs☐
7Why Actor-Critic merges the two lines☐
8Why we need advantage / baselines☐
9Why DQN needs replay + target networks☐
10What Double/Dueling/Rainbow each solve☐
11REINFORCE's variance problem and its fixes☐
12PPO clipped objective: formula and intuition☐
13PPO vs TRPO☐
14What GAE is, what λ tunes☐
15SAC's maximum-entropy objective and automatic temperature☐
16TD3's three tricks☐
17Exploration vs exploitation: ε-greedy / UCB / Thompson sampling compared☐
18How exploration works in deep RL (entropy, NoisyNets, RND)☐
19Multi-armed bandits vs full RL (what's missing in a bandit)☐
20Reward shaping and the intuition behind the potential-based theorem☐
21Reward-hacking cases and defenses☐
22Sparse-reward countermeasures (curricula, HER, intrinsic rewards)☐
23The three-stage RLHF pipeline☐
24Reward models and Bradley–Terry☐
25Why the KL penalty in PPO is necessary☐
26Reward over-optimization and Goodhart☐
27The DPO formula and its trade-offs vs PPO☐
28The OOD problem and value overestimation in offline RL☐
29Solution routes when the model is known vs unknown☐
30Modeling an MDP live from a business scenario (Q10/Q11 frameworks)☐

Self-test rules

Checking every box doesn't mean you know it — you only know it if you can survive the follow-ups. For every question you check "yes," find someone (or record yourself) and answer that question's "likely follow-ups" out loud too. Can't answer the follow-ups? Send the question back to the Knowledge Map and re-mark it "half-know."

Further Reading ​

References ​

  • Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd ed. — the authoritative source for theory questions, free online: http://incompleteideas.net/book/the-book-2nd.html
  • Schulman, J., et al. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347 — the original PPO clipped objective.
  • Mnih, V., et al. (2015). Human-level control through deep reinforcement learning. Nature 518. arXiv:1509.06461 — the original DQN paper.
  • Haarnoja, T., et al. (2018). Soft Actor-Critic. arXiv:1801.01290 — the original SAC paper.
  • Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback (InstructGPT). arXiv:2203.02155 — the original three-stage RLHF paper.
  • Rafailov, R., et al. (2023). Direct Preference Optimization. arXiv:2305.18290 — the original DPO paper.
  • Christiano, P., et al. (2017). Deep reinforcement learning from human preferences. arXiv:1706.03741 — the early form of RLHF.
  • OpenAI Spinning Up in Deep RL (algorithms matched with code): https://spinningup.openai.com