Appearance
RLHF and Alignment with Human Feedback
In a nutshell: this page is all about RLHF (reinforcement learning from human feedback) — the key technique that turns a large language model from "able to continue a sentence" into "able to talk like a human": the three-stage SFT → reward model → PPO pipeline, why the KL constraint is indispensable, how reward overoptimization happens, and how DPO collapses it all into a single step. By the end you'll be able to draw the full RLHF data flow, explain the role of every component, and articulate how it compares with classic RL.
1. The alignment problem: why "pretraining + SFT" is not enough
1.1 The problem: capability is not conversation
A pretrained model's job is to "predict the next token" — it knows how to continue a piece of text, but not what a good reply actually looks like. Supervised fine-tuning (SFT) teaches it to "follow instructions" using human-written (instruction, desired response) pairs.
But SFT has three hard limitations:
| Limitation of SFT | Explanation |
|---|---|
| Imitation ceiling | It can only imitate the responses in the SFT data — it never rises above the quality of the examples |
| Data is expensive | High-quality, human-written instruction–response pairs are extremely costly |
| Cannot express "preference" | The data only says "this response is good", never "of these two responses, which is better, and by how much" |
The last point matters most: real human preferences are relative, multi-dimensional, and hard to capture in a single example. Ask for "a paper abstract" and two AI responses may both be decent while one is clearly better. SFT has no way to encode this kind of "relatively better".
1.2 The core insight of RLHF
Instead of giving the model the "correct answer", give it a "comparison signal": which of two responses is better. Comparison data is far cheaper to produce than written demonstrations, and it encodes human preferences natively.
RLHF (the InstructGPT/ChatGPT recipe, Ouyang et al., 2022) makes alignment a matter of "optimizing with RL a reward learned from human preferences". This is why it earns a place in an RL handbook: RLHF is the canonical RL instance where the reward is defined by human preference — the environment is the token-generation process, and the reward is the reward model's score.
2. The three-stage pipeline: SFT → reward model → PPO
text
RLHF: three stages (the InstructGPT recipe)
┌────────────────────────────────────────────────────────────┐
│ Stage 1: SFT (supervised fine-tuning) │
│ data: (instruction, human-written ideal response) │
│ output: π_SFT — a model that "follows instructions" │
│ │
│ Stage 2: Reward model (RM) │
│ data: (instruction, response A, response B, │
│ human label "which is better") │
│ output: r_φ(x, y) — a model that scores responses │
│ │
│ Stage 3: PPO fine-tuning │
│ models: policy π_θ (initialized from π_SFT) │
│ value net V_ψ + reward model r_φ + ref π_ref │
│ objective: max E[r_φ] − β·KL(π_θ ‖ π_ref) │
│ output: the aligned model π_θ │
└────────────────────────────────────────────────────────────┘2.1 Stage 1: SFT (the starting point of behavior cloning)
Supervised fine-tuning on a few thousand to tens of thousands of human-written (instruction, high-quality response) pairs. The purpose of this stage is to "pull the model back onto the rails of instruction following"; it is the behavioral foundation for everything that follows. The quality of the SFT data directly sets the ceiling — and, as covered in Policy Gradient Methods, it is behavior cloning applied at scale, with the limitations discussed above.
2.2 Stage 2: the reward model (turning preferences into scores)
The reward model $r_\phi(x, y)$ takes (instruction x, response y) and outputs a scalar score. Its training data consists of preference pairs:
instruction x: "Explain quantum entanglement in one sentence"
response A: "Quantum entanglement is..." ← the human labeler picks A as better
response B: "Entanglement is entanglement"The Bradley–Terry model writes the "preference probability" as a sigmoid of the reward difference:
$$ P(A \succ B \mid x) = \frac{\exp(r_\phi(x, A))}{\exp(r_\phi(x, A)) + \exp(r_\phi(x, B))} = \sigma\big(r_\phi(x,A) - r_\phi(x,B)\big) $$
Training objective: maximize the log-likelihood of the human-labeled preference pairs:
$$ \mathcal{L}{RM} = -\mathbb{E}{(x, A, B) \sim D}\left[ \log \sigma\big(r_\phi(x,A) - r_\phi(x,B)\big) \right] $$
The intuition: a reward model is a "learned scorer" — it distills human "which is better" judgments into a differentiable function that scores any response. That function then becomes the reward source for RL optimization.
Why preference pairs instead of direct scores
Asking people to score responses from 1 to 5 gives inconsistent scales (one annotator's 4 is another's 3) and is mentally taxing. Binary choices are cheaper and more consistent — relative judgments are far more reliable than absolute ones. This is the same idea as the "preference learning" section in Reward Engineering.
2.3 Stage 3: PPO fine-tuning (RL inside "prompt sampling")
This step ports classic PPO onto a language model:
- Environment = a prompt x, plus the model's token-generation process;
- Action = the generated response y (a sequence of tokens);
- Reward = the reward model's score $r_\phi(x,y)$ (sparse: one score per response);
- Policy = the LLM itself, $\pi_\theta$.
The optimization objective (the core formula of RLHF):
$$ \max_\theta \underbrace{\mathbb{E}{x \sim \mathcal{D}, y \sim \pi\theta}[r_\phi(x,y)]}{\text{make the reward high}} - \underbrace{\beta , D\big( \pi_\theta(y|x) ,|, \pi_{ref}(y|x) \big)}_{\text{don't drift too far}} $$
Here $\pi_{ref}$ is the reference model (usually π_SFT or the previous version of the policy), and β is the KL-penalty coefficient.
2.4 Why the reference model and KL penalty are mandatory
In one sentence: the reward model is only "locally trustworthy" — it has judgment only near "normal responses". Once the policy generates text far from the training distribution (gibberish, broken syntax, extreme wording), the reward model's scores are completely unreliable (an out-of-distribution problem; see Offline RL). The KL penalty does the following:
| Role of KL | Explanation |
|---|---|
| Caps the exploration radius | The policy is not allowed to drift too far from the reference model (see Section 8 of Exploration and Exploitation) |
| Prevents reward-model hallucination | Stops the policy from optimizing into regions where the reward model mistakenly gives high scores |
| Preserves language quality | Prevents collapse into keyword-stuffing "reward farming" |
| Provides an entropy-like constraint | Similar to entropy regularization in classic RL, keeping outputs diverse |
What happens if you remove the KL term
In training you will soon see "reward soaring, language collapsing" — the policy discovers a hole in the reward model (for example, endlessly repeating "quantum entanglement quantum entanglement quantum entanglement..." to harvest points). The KL penalty is the non-negotiable seatbelt of RLHF engineering. Remove it and you get the reward hacking from Reward Engineering all over again — this time on a massive model.
3. The cast of models on stage in RLHF training
During PPO, four models are actually running at once — the heaviest part of RLHF engineering:
| Model | Role | Needs gradients? |
|---|---|---|
| Policy model π_θ (actor) | Generates responses; gets optimized | Yes |
| Value model V_ψ (critic) | Estimates the state value of the current generation sequence (for the advantage) | Yes |
| Reward model r_φ | Scores responses | No (frozen, inference only) |
| Reference model π_ref | The reference for computing the KL penalty | No (frozen, inference only) |
The training loop (the PPO-ptx variant from InstructGPT):
text
loop:
1. Sample a prompt x ~ D
2. Generate a response y with π_θ (one rollout)
3. reward = r_φ(x,y) + KL penalty (against π_ref)
4. Compute the advantage (GAE, needs V_ψ)
5. Update π_θ (clipped objective + entropy bonus)
6. Update V_ψ (TD regression)
7. Optional: mix in 10% SFT data to prevent "catastrophic forgetting"
(the PPO-ptx trick)RLHF through the eyes of an RL engineer
The PPO stage of RLHF is a standard on-policy actor-critic; the only special feature is that the reward comes from a learned reward model, and each episode is a single response — one sparse reward. GAE is nearly vacuous here (there are no intermediate-step rewards), and the advantage is provided mostly by the "step-by-step accumulation of the KL penalty". This is why PPO implementations for RLHF are often trimmed down to something quite lightweight.
4. The alignment tax and reward overoptimization (Goodhart's law)
4.1 The alignment tax
Alignment tax: the alignment process (RLHF) costs the model points on standard capability benchmarks (such as MMLU or code evaluations). "Pleasing human preferences" and "general capability" do not fully align — over-catering makes the model "more polite but more mediocre". The InstructGPT paper reports this capability regression explicitly and mitigates it by "mixing in SFT data".
4.2 Reward overoptimization: higher reward ≠ a better model
This is the most critical empirical phenomenon in RLHF: the longer training runs, the higher the reward-model score climbs — but human ratings of the actual output rise first and then fall.
text
quality ▲
│ ╱‾‾‾‾╲
│ ╱/ ╲ ← human rating (true quality)
│ ╱/
│ ╱
├──────╲──────────────► training steps
│ ╲ reward model score (keeps climbing)
│ ╲
└──────────────────────────────►
true quality starts declining past a point,
while the reward model still "thinks" things are improving.This is Goodhart's law: "When a measure becomes a target, it ceases to be a good measure." The reward model is a "proxy for human preference"; optimizing it to the extreme means optimizing a proxy that has drifted away from the real objective.
4.3 Engineering countermeasures
| Technique | Mechanism |
|---|---|
| KL constraint | Cap the exploration radius; don't let the policy leave the reward model's trustworthy region |
| Early stopping | Monitor the reward-vs-true-quality divergence curve and stop before the turning point |
| Reward model ensembles / calibration | Have multiple RMs vote; calibrate with human spot checks |
| Mixed objectives | PPO-ptx mixes in SFT data to prevent forgetting |
| Preference-label spot checks | Periodically evaluate the final policy with humans instead of trusting RM scores alone |
5. DPO: treating preferences as the objective directly (bypassing the reward model)
5.1 Motivation: RLHF is heavy engineering
RLHF requires maintaining four models plus a PPO training loop; the engineering is complex and training is unstable (the reward model, KL, and GAE all need tuning). DPO (Direct Preference Optimization, Rafailov et al., 2023) proved that the reward-model stage can be eliminated mathematically.
5.2 Core idea: a closed-form translation of "preferences" into a policy objective
The key insight of DPO: the RLHF optimization problem (maximize reward − KL penalty) has a closed-form optimal solution:
$$ \pi^*(y|x) \propto \pi_{ref}(y|x) , \exp\left( \frac{1}{\beta} r_\phi(x,y) \right) $$
Invert this to recover the reward: $r(x,y) = \beta \log\frac{\pi^*(y|x)}{\pi_{ref}(y|x)} + \text{const}$. Substituting "reward = the policy-to-reference ratio" back into the Bradley–Terry preference loss makes the reward model disappear entirely, yielding the DPO loss:
$$ \mathcal{L}{DPO}(\theta) = -\mathbb{E}{(x, A, B) \sim D}\left[ \log \sigma\left( \beta \log \frac{\pi_\theta(A|x)}{\pi_{ref}(A|x)} - \beta \log \frac{\pi_\theta(B|x)}{\pi_{ref}(B|x)} \right) \right] $$
The intuition: DPO trains the policy directly on preference pairs — make the policy more likely to produce the "preferred A" and less likely to produce the "rejected B", with the KL ratio controlling the size of the change (β is the KL strength). No reward model, no PPO.
5.3 RLHF vs. DPO
| Dimension | RLHF (PPO route) | DPO |
|---|---|---|
| Components | SFT + RM + PPO (4 models) | SFT + preference pairs (2 models) |
| Training stability | Poor (PPO is finicky) | Good (simple classification-style training) |
| Hyperparameters | Many (β, KL, clip…) | Few (β) |
| Scalability | Heavy, expensive | Light, cheap |
| Robustness to reward off-distribution | KL as the safety net | Also constrained by the β ratio |
| Online sampling | Yes (RL-style) | Mostly offline (on-policy sampling variants exist) |
| Representative models | InstructGPT, ChatGPT, Llama 2-Chat | Zephyr and the DPO family |
How to choose today
- With engineering resources, chasing the ceiling: RLHF (especially online RLHF — see below); better results and easy to iterate;
- Fast alignment, open-source replication: the DPO family (Zephyr et al.) — up and running within a day;
- The academic and industry consensus: the DPO simplification is the direction of travel, but whether you need "PPO-grade RL" depends on your budget and quality bar. Frontier variants include IPO, KTO, ORPO, and more (see Frontier Advances).
6. RLHF vs. classic RL
| Dimension | Classic RL | RLHF |
|---|---|---|
| Environment | Grids / robots / games | The token-generation process |
| State s | Environment state | Prompt + tokens generated so far |
| Action a | Joint torques / directions | The next token |
| Transition P | Physics / game rules | The model's own autoregressive sampling |
| Reward R | Hand-designed (see Reward Engineering) | Scores from a learned reward model |
| Sparsity | Often sparse | Extremely sparse (one score per response) |
| Exploration | ε / entropy / intrinsic rewards | KL constraint + sampling temperature |
| Data source | Online trial and error | Human preference labels + self-sampling |
| Evaluation | Return curves | Human evaluation + RM scores + capability benchmarks |
A curious inversion: in classic RL we hand-write the reward (because the task is quantifiable); in RLHF we learn the reward (because preferences resist being put into words). The two converge on the unified framework of "optimizing return" — RLHF is a direct instance of the classic RL framework applied to language generation.
7. Limitations and frontiers
7.1 Inherent limitations of RLHF
| Limitation | Explanation |
|---|---|
| Preference-label bias | Annotators are subjective; position and order biases exist |
| Reward overoptimization | Goodhart's law; needs early stopping + KL (see Section 4) |
| Aligns only to "average human preference" | Diverse users and minority needs get averaged away |
| Optimizes for "looking good" | No guarantee of factual correctness (hallucination still needs external fixes such as retrieval augmentation) |
| High engineering cost | 4 models + PPO + distributed RL training |
7.2 Frontier directions
- RLAIF (RL from AI feedback): use another LLM as the "labeler" to generate preference pairs in place of human annotators, greatly improving scalability (Bai et al., 2022, Constitutional AI);
- Online RLHF: alternate reward-model and policy training (e.g., self-play-style preference sampling), mitigating the staleness of offline preference data;
- RLHF for reasoning (RLVR): replace preferences with "verifiable rewards" (does the code pass the tests, is the math answer right) — when the rules can verify, no human feedback is needed and you do straight RL (the DeepSeek-R1 route); see Frontier Advances;
- The whole DPO variant family: IPO, KTO, SimPO, ORPO, and more, each trading off simplicity against robustness.
7.3 Historical lineage
RLHF did not appear out of thin air: it inherits the actor-critic of Value-Based Learning, the PPO of Policy Gradient Methods, and the preference learning of Reward Engineering. For the full history and the classic papers, see A Brief History of RL and Classic Papers, Close Reading (the Ouyang InstructGPT entry).
8. Key takeaways for engineers
- SFT is the foundation: neither the reward model nor PPO can rescue a model trained on poor SFT data;
- Reward-model quality sets the ceiling: effort spent on preference-data cleaning and labeling consistency pays off far more than tuning PPO hyperparameters;
- The KL penalty is non-negotiable and β must be tuned: β too large → you revert to SFT (no learning); β too small → reward hacking;
- Monitor the divergence curve: reward vs. human spot-check quality, and stop before the turning point;
- DPO first, RLHF second: on a limited budget, get a baseline with DPO first, then invest in RLHF to push the ceiling.
For practical details (open-source replication, three-model training, cost accounting), see LLM Alignment in Practice: RLHF; for the conceptual discussion of reward design and hacking, see Reward Engineering.
Further Reading
- LLM Alignment in Practice: RLHF — the full engineering picture of InstructGPT/ChatGPT and open-source replication
- Policy Gradient Methods — the PPO clipped objective and GAE, the optimizer foundation of RLHF
- Reward Engineering — the conceptual source of preference learning and reward hacking
- Classic Papers, Close Reading — a paragraph-by-paragraph close reading of the InstructGPT paper
- Frontier Advances — the latest on the DPO family, RLAIF, online RLHF, and RLVR
- A Brief History of RL — seventy years of coordinates, from Bellman to RLHF
References
- Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training language models to follow instructions with human feedback. NeurIPS. arXiv:2203.02155 (InstructGPT — the standard citation for RLHF)
- Rafailov, R., Sharma, A., Mitchell, E., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS. arXiv:2305.18290 (DPO)
- Bai, Y., Jones, A., Ndousse, K., et al. (2022). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv:2204.05862 (Constitutional AI / RLAIF)
- Stiennon, N., Ouyang, L., Wu, J., et al. (2020). Learning to summarize with human feedback. NeurIPS. arXiv:2009.01325 (the precursor work to RLHF)
- Ziegler, D. M., Stiennon, N., Wu, J., et al. (2019). Fine-Tuning Language Models from Human Preferences. arXiv:1909.08593 (an early version of RLHF)
- Christiano, P. F., Leike, J., Brown, T. B., et al. (2017). Deep Reinforcement Learning from Human Preferences. NeurIPS. arXiv:1706.03741 (where the RLHF concept was introduced)