Appearance
LLM Alignment: RLHF in Practice
In a nutshell: this page explains, from an RL engineer's perspective, what RLHF (reinforcement learning from human feedback) actually is — how the three-stage SFT → reward model → PPO pipeline works, why the reference model and the KL penalty are necessary, the true costs of reward overoptimization and the "alignment tax," and how DPO compresses all of this into a single idea. It is the full worked example accompanying the RLHF and Human-Feedback Alignment concept page.
1. Background: Pretrained Models "Continue Text" — They Don't "Converse"
1.1 Where the Problem Comes From
By 2020, GPT-3 could "continue" any text — but it had three cardinal sins:
| Symptom | How it shows up | Root cause |
|---|---|---|
| Won't follow instructions | Ask it to "help me write an email" and it may just continue the sentence instead of actually answering | The pretraining objective is "next-token probability," not "get the task done" |
| Harmful / biased output | It produces discriminatory, offensive, or false content | It learned the bad parts of internet corpora |
| Confidently wrong | It fabricates facts and delivers wrong answers with a straight face | There is no "right/wrong" signal — only "does it sound like human text" |
The fundamental tension: pretraining optimizes for "more like the corpus," while a product needs "more helpful, more truthful, safer" — two different objectives. The latter is what the field calls alignment.
1.2 Why "Just Train Another SFT" Isn't Enough
Supervised fine-tuning (SFT) does help — fine-tune on human-written "instruction–answer" pairs and the model learns what an answer looks like. But SFT has two ceilings:
- Data volume: high-quality human-labeled instruction–answer pairs are expensive, capping out in the tens of thousands — nowhere near covering the behavior space;
- The imitation ceiling: SFT only imitates the annotators' "demonstrations" — it never beats them and can't handle "edges the demonstrations didn't cover." This is the "limits of behavior cloning" from the multi-armed bandits page, magnified to the scale of large models.
So in the 2022 InstructGPT paper, OpenAI delivered the full answer: turn "human preference" into an optimizable reward signal, then optimize it with reinforcement learning (PPO) — that is RLHF.
2. The Three-Stage Pipeline: SFT → Reward Model → PPO
text
Stage 1 SFT: supervised fine-tune Stage 2 RM: reward model Stage 3 RL: PPO optimization
┌──────────────────────────────┐ ┌──────────────────────────────┐ ┌──────────────────────────────┐
│ Humans write │ │ Model generates │ │ RM score as the reward; │
│ instruction-answer pairs; │ │ multiple answers; humans │ │ PPO updates the policy; │
│ fine-tune w/ cross-entropy │──► │ rank/pair them; train │──► │ KL penalty against the │
│ (learn the answer shape) │ │ a scorer RM (predicts │ │ frozen reference model │
│ tens of thousands of samples │ │ human preference) │ │ (learn better answers) │
└──────────────────────────────┘ └──────────────────────────────┘ └──────────────────────────────┘| Stage | Input | Training objective | Resources required |
|---|---|---|---|
| SFT | Human instruction–answer pairs | Maximize the log probability of the answers | Tens of thousands of labeled samples |
| RM (reward model) | Multiple answers to the same prompt + human preferences | Give "more preferred" answers higher scores | Hundreds of thousands of preference pairs |
| PPO | Prompt pool + RM scores | Maximize "reward − KL penalty" | Substantial inference compute |
TIP
In RLHF, the "environment" is the language model itself: the action is "generate a token / an answer," the state is "the sequence of tokens generated so far," and the reward is the RM's final score. This maps one-to-one onto the "RLHF vs. classic RL" table on the RLHF and Human-Feedback Alignment concept page — it is an episodic, one-shot sequential decision problem with no intermediate rewards, only a terminal score.
3. Training the Reward Model: The Bradley-Terry Intuition
3.1 Why Humans Can't Be the Reward Directly
If a human had to judge "is this right?" after every generated token, they'd burn out — and their labels would be inconsistent anyway. RLHF's solution is to train a surrogate reward function (RM): have humans compare answers in batches, then learn a scoring model from those comparisons.
3.2 Preference Data and Bradley-Terry
Given the same prompt x, the model generates two answers — y_w (the better one) and y_l (the worse one) — and a human labels which is better. The output of the reward model r_θ(x, y) is interpreted as the probability that a human prefers a given answer, via the Bradley-Terry model:
text
P(y_w preferred over y_l) = σ( r_θ(x, y_w) − r_θ(x, y_l) )
σ = sigmoid function
Loss: −log σ( r_θ(x, y_w) − r_θ(x, y_l) )
Intuition: push the score of "the answer humans like" well above "the answer they don't"The RM is typically a sequence model that emits a scalar score at the final token. InstructGPT trained it on roughly 33,000 prompts, each with multiple answers, yielding preference data at that scale.
3.3 Key Details of Preference Data
- Within-prompt comparisons: comparisons must be between "two answers to the same prompt"; cross-prompt comparisons are meaningless;
- Labeler consistency: measure agreement among annotators on the same answer pair (e.g., Cohen's κ);
- Ranking vs. scoring: "ranking" is a more robust annotation format than "absolute scores" — humans are better at ordering things than scoring them.
This whole recipe of "learning a reward from pairwise preferences" is exactly the deployed form of the "preference learning: deriving rewards from human preferences" section on the reward engineering page.
3.4 Reward Model Generalization and Calibration
The RM is what PPO directly optimizes, so its quality sets the ceiling for RLHF. Three engineering points:
| Point | Practice | Pitfall |
|---|---|---|
| Cover the training distribution | Preference prompts must cover "the prompts the policy will see" | Train only on safe prompts and the RM scores dangerous ones wildly |
| Reward mean calibration | Re-score after each new model generation; track RM score drift | Systematically inflated RM scores → misjudged overoptimization |
| Stay in sync with the policy | When the policy changes, re-validate the RM's "preference judgments" | An old RM evaluating a new policy produces distorted assessments |
A pragmatic check: during PPO training, periodically measure the RM's accuracy on a batch of human-labeled answers — if its accuracy on that data drops below 70%, something is wrong with the preference data or the RM itself; fix the RM before training the policy further. This is exactly the "independent evaluation set" idea from the evaluation and benchmarks page.
4. The PPO Stage: Four Models on Stage at Once — Actor / Critic / Reference / RM
4.1 The Roles of the Four Models
PPO fine-tuning is the heaviest engineering step in RLHF — four models run at the same time:
text
actor π_θ : the model being trained; produces answers (the policy)
reference π_ref : the frozen SFT model; measures "how far the policy has drifted" (KL baseline)
reward r_φ : the RM trained in stage 2; scores answers
critic V_ψ : value network; estimates "expected reward−KL from the current state" (needed by PPO)
────────────────────────────────────────────────────────────
Training loop (per prompt):
1. The actor samples an answer (a complete episode)
2. The RM scores the whole answer: r_φ(x, y)
3. Per-token KL penalty: r̂_t = r_φ(x,y) only at the final token − β·KL(π_θ ‖ π_ref)
4. Update the actor with PPO; update the critic with GAE4.2 Why the Reference Model and KL Penalty Are Needed
Maximize the RM score directly, and the model will find the RM's loopholes (reward hacking) — saying whatever fools the RM into awarding high points. The reference model π_ref (usually the SFT model) provides an "anchor of reasonable language," penalizing the policy per token by its KL distance from that anchor:
text
KL penalty term: −β·D_KL( π_θ(·|s_t) ‖ π_ref(·|s_t) )
Meaning: every time the policy "drifts from the reference model's distribution," points are deducted.
Effect: the model is allowed to improve, but not to lose the ability to "talk normally" in exchange for reward.The intuition: this is the answer to the "why KL?" section on the RLHF concept page — the KL penalty is RLHF's "exploration guardrail": it prevents the policy from degenerating entirely (babbling nonsense to farm reward) and from distribution collapse (only ever saying one thing). Seen from the exploration vs. exploitation angle, the KL constraint is RLHF's distinctive form of exploration: exploring within a neighborhood of the reference distribution.
4.3 Hyperparameters and Training Realities
| Item | Typical values (InstructGPT / open-source reproductions) | Notes |
|---|---|---|
| KL coefficient β | On the order of 0.01–0.1 | Larger = more conservative |
| PPO clip range | 0.2 | Same as classic PPO |
| Learning rate | Separate for actor/critic, on the order of 1e-6 to 1e-5 | Fine-tuning large models requires baby steps |
| Batch | Hundreds to thousands of prompts | One answer sampled per prompt |
| Training steps | A few thousand | Alignment doesn't need many steps; overtraining degrades |
Alignment training is "lightweight" in scale: the InstructGPT paper showed that 1.3B InstructGPT beat the 175B unaligned GPT-3 in human evaluations — alignment makes "small models more useful," rather than relying on parameter count.
4.4 Implementation Notes for PPO on LLMs
PPO in RLHF differs from PPO in robotics or games in three implementation details that are easy to trip over:
| Implementation point | What it means | Common pitfall |
|---|---|---|
| Reward only at the final token | The RM score for the whole answer is added only to the last token; all earlier tokens get zero reward | Spreading the RM score across every token breaks the semantics |
| Per-token KL | The KL penalty is computed distribution-wise at each token position, not once for the whole answer | Skipping per-token KL leads to severe overoptimization |
| Reference model stays frozen | π_ref is never updated; it is the "anchor" | Accidentally training π_ref too disables the KL constraint |
| Critic mirrors the actor | The value network usually copies the actor architecture (shared backbone or standalone) | A critic that is too small estimates GAE poorly |
WARNING
One more point that even RL veterans tend to miss: RLHF "episodes" are extremely short (an answer of a few dozen to a few hundred tokens), so the semantics of GAE's λ differ from those in long-horizon environments. In practice, many open-source reproductions skip fine-tuning λ entirely — the default works well enough. But if you're doing RLHF on long-answer tasks (e.g., long-document summarization), revisit λ and reward normalization.
5. Case Studies, Data, and Results
5.1 Human Evaluation of InstructGPT
| Metric | Result |
|---|---|
| 1.3B InstructGPT vs. 175B GPT-3 | Humans preferred InstructGPT's outputs (by a significant majority) |
| Helpfulness | Far ahead of the unaligned model |
| Truthfulness | Hallucinations reduced, though not eliminated |
| Toxicity | Mitigated on adversarial prompts |
5.2 LLaMA-2-chat: A Complete Open-Source Loop
Meta released LLaMA-2 in 2023, and LLaMA-2-chat disclosed complete RLHF details:
- Multi-round RLHF: SFT → two or more "reward model + PPO" iterations; in each round, preference labels are re-collected on samples from the previous round's model, the RM is updated, and the policy is retrained;
- A family of reward models: separate "helpfulness RM" and "safety RM" trained on different prompt distributions;
- Rejection sampling fine-tuning: filter sampled answers through the RM, then SFT directly on the high-reward samples before PPO — an "offline distillation" trick that reduces PPO's exploration cost;
- The alignment tax, documented: Meta reported that after RLHF, capability stayed flat or dipped slightly on most tasks while safety and helpfulness improved — public evidence for the "alignment tax."
5.3 Zephyr: Simplified Open-Source Alignment with DPO
HuggingFace's Zephyr-7B (Tunstall et al., 2023) proved the engineering value of DPO (Direct Preference Optimization): no RM training, no PPO — just one "closed-form" fine-tune on preference pairs — and a 7B model matched or beat stronger open-source models of the day on benchmarks like MT-Bench.
5.4 Dataset Composition and the Alignment Cost Ledger
RLHF's true cost is routinely underestimated; break it down to see:
| Stage | Main cost | Rule-of-thumb share |
|---|---|---|
| Preference data labeling | Humans labeling hundreds of thousands of preference pairs | High (labor) |
| RM training | Moderate compute | Low |
| PPO sampling and training | Multiple samples + four model forward passes | High (compute) |
| Evaluation and regression | Independent evaluation set + human spot checks | Medium |
InstructGPT's approach was to sample prompts from real API user requests (rather than hand-writing prompts), covering the true distribution; each prompt got multiple answers, which annotators compared pairwise, producing a dataset of ~33,000 prompts and hundreds of thousands of preference pairs. This lesson is worth copying: for preference data, "distribution coverage" matters more than "raw volume" — sampling from real usage is far more efficient than inventing prompts in a vacuum.
WARNING
Alignment cost can't be judged by a single training run: as models iterate, preference data goes stale (it labels outputs of older models) and must be continuously re-sampled and re-labeled. The cost of maintaining a "preference data pipeline" over the long run is often several times that of any single training run. This is the hidden bill every alignment team runs into.
6. The Alignment Tax and Reward Overoptimization: Goodhart's Law
6.1 Two Real Costs
Alignment tax: alignment training can make the model worse at tasks it used to be good at (math, code) — because preference data favors "likable" over "correct." LLaMA-2's report of slightly lower scores on some benchmarks after RLHF is exactly this cost.
Reward overoptimization: as PPO steps pile up, the RM score keeps climbing while true human evaluation scores start falling — the model is "farming the RM's points" instead of getting better. This is Goodhart's law ("when a measure becomes a target, it ceases to be a good measure") in its standard RLHF form:
text
RM score
↗↗↗
training steps ────►
↘↘
true human-preference score (rises, then falls)Countermeasures:
- KL penalty: smaller β is not better — too small and you get overoptimization;
- Early stopping: decide when to stop using human evaluation or a held-out RM validation set — don't keep training just because the RM score is still rising;
- DPO's "implicit KL": DPO bakes the KL constraint into the objective itself, making it inherently less prone to overoptimization than PPO (though not immune).
WARNING
The most common RLHF failure in industry is "the RM score is still going up, so let's train a few more rounds." The RM score is a proxy metric, not the real objective; the real objective is "humans find it useful." The right approach is an evaluation set + human spot checks independent of the training RM, treating "RM score plateauing or dropping" together with "human evaluation dropping" as the early-stopping signal. This corresponds exactly to the "don't make decisions with a single metric" discipline on the evaluation and benchmarks page.
7. DPO: Compressing RLHF into One Sentence
7.1 The Core Idea of DPO
The key insight of Rafailov et al. (2023): the optimal solution of "reward maximization + KL constraint" in PPO can be written in closed form in terms of the reward function r(x, y). This lets you eliminate PPO and the RM, turning "preference" directly into a classification loss:
text
Optimal policy for PPO's objective (maximize reward − β·KL):
π*(y|x) ∝ π_ref(y|x) · exp( r(x,y) / β )
Solve for the reward in reverse:
r(x,y) = β·log[ π_θ(y|x) / π_ref(y|x) ] + constant
Substitute r back into the Bradley-Terry preference loss -> the DPO loss:
L = −log σ( β·log[ π_θ(y_w|x)/π_ref(y_w|x) ] − β·log[ π_θ(y_l|x)/π_ref(y_l|x) ] )Intuition: DPO no longer trains a reward model explicitly; it directly requires the policy to raise the relative probability of y_w and lower that of y_l (relative to the reference model). It keeps RLHF's "KL constraint + preference objective" while eliminating all the engineering complexity of RM training and PPO.
7.2 PPO vs. DPO
| Dimension | PPO-RLHF | DPO |
|---|---|---|
| Reward model | Must be trained separately | Not needed |
| Training stability | Sensitive (four models, many hyperparameters) | Simple and stable |
| KL constraint | Explicit β coefficient | Implicit (reference model in the objective) |
| Sample efficiency | Requires large amounts of online sampling | Offline preference data suffices |
| Scalability | Strong at online feedback / iterative alignment | Excels offline; online variants emerging |
| Open-source ecosystem | LLaMA-2-style RLHF pipelines | Widely adopted by Zephyr, Mistral-family models |
7.3 Which One, When
- Want "good chat experience, simple engineering": DPO is enough — Zephyr is the proof;
- Need "online iteration" (the model generates, preference data is continuously refreshed): online RLHF in the PPO family is the better fit;
- Research frontier: the two are converging (online variants of DPO, GRPO, etc.) — see Frontier Advances.
7.4 The Full Variant Landscape: PPO / DPO / IPO / KTO / GRPO
After 2023, alignment algorithms exploded — one table to see the mainstream variants:
| Algorithm | Core idea | Needs RM? | Needs reference model? | Characteristics |
|---|---|---|---|---|
| PPO-RLHF | RM scoring + KL penalty + policy optimization | Yes | Yes | Most general; heavy engineering |
| DPO | Closed-form solution eliminates the RM and sampling | No | Yes | Simple, stable, offline |
| IPO | Fixes DPO's overfitting (uses preference pairs directly) | No | Yes | Steadier, simpler sampling |
| KTO | Only needs "good/bad" labels, no pairing | No | Yes | Lower data bar |
| GRPO | Relative rewards within a group, drops the critic (adopted by DeepSeek) | No | Yes | Saves memory; suits reasoning RL |
Quick selection guide: paired preference data + simplicity → DPO; only good/bad labels → KTO; large-scale reasoning RL (math/code) → the GRPO route; online iteration → the PPO family.
8. The Frontier: RLAIF, Online RLHF, and Reasoning Alignment
8.1 RLAIF: Replacing Human Feedback with AI Feedback
RLAIF (RL from AI feedback): instead of having humans label preferences, a strong LLM (say, GPT-4) acts as the "judge," generating preference labels that feed the same training pipeline. Early work from DeepMind and the public study of Lee et al. (2023) showed that AI feedback can approach human-feedback quality on tasks like harmlessness, at far better cost and scalability than human labeling. It turns RLHF from "labor-intensive" into "scalable."
8.2 Online RLHF and Reasoning Alignment (RLVR)
The hottest direction in 2024–2025 is using RL directly to improve LLM reasoning:
- RLVR (RL with verifiable rewards): for tasks where answers can be verified automatically (math, code), use a rule-based verifier (correct/incorrect) as the reward, bypassing the RM entirely — the reward is objective, with no dependence on human preference;
- DeepSeek-R1 and friends: large-scale RL teaches the model "chain-of-thought self-reflection," yielding huge gains in math/code reasoning — with no SFT cold-start stage (pure RL straight from the base model);
- This direction expands RLHF from "aligning human preferences" to "optimizing verifiable capabilities," and it is the star of the "RL for reasoning" section on the frontier advances page.
8.3 Convergence with Agent RL
When an LLM is used as an agent (calling tools, browsing the web, executing multi-step tasks), the "reward" shifts from "human preference" to "did the task succeed" (was the tool call correct? did the web action achieve the goal?) — the RLHF framework carries over unchanged, with only the reward source swapped. That's the subject of the Agents and Dialogue Systems page.
9. Key Takeaways for RL Engineers
9.1 RLHF vs. Classic RL
| Classic RL | RLHF | Correspondence |
|---|---|---|
| Environment | The language model itself (no real environment) | Environment = policy |
| Reward | Environment signal | RM score (proxy reward) |
| Policy | π(a|s) | Language model π(y|x) |
| Exploration | ε-greedy, entropy | KL to the reference model (exploration guardrail) |
| Sparse rewards | Common | RLHF has a single terminal score (sparsest possible) |
| Value estimation | Q/V networks | Critic network |
| Engineering focus | Environment and data | Preference data and training stability |
9.2 What RL Engineers Can Bring Over
- All the PPO machinery transfers directly: GAE, clipping, value clipping, gradient clipping;
- Evaluation discipline: multiple seeds, independent evaluation sets, early stopping — all just as valid in RLHF;
- Reward-engineering instincts: the RM is "the reward function," and the detection methods for reward hacking (see reward engineering) are called "reward overoptimization" in the LLM world — same thing.
9.3 Common Pitfalls (From an RL Perspective)
| Pitfall | Symptom | Countermeasure |
|---|---|---|
| KL coefficient tuned haphazardly | Outputs become perfunctory or unhinged | Run ablations on β; don't settle for the default |
| RM out of sync with the actor | Inflated RM scores | Recalibrate the RM regularly on fresh model outputs |
| Overtraining | Human evaluations drop | Independent evaluation set + early stopping |
| Dirty preference data | The RM learns the wrong thing | Label-consistency checks, prompt diversity |
| Treating the proxy metric as the goal | Users unhappy after launch | Close the loop with human evaluation / live metrics |
9.4 Five Must-Know RLHF Interview Questions (with Answer Points)
RLHF is a high-frequency topic in RL job interviews. Here are the five inevitable questions, with answer points:
| Question | Answer points |
|---|---|
| Why does RLHF need three stages — why not PPO directly? | You need a reward function first (the RM); the RM needs preference data; PPO with no signal goes nowhere |
| Who gets the KL penalty, and why? | Applied between the policy and the reference model, to prevent reward hacking and distribution collapse |
| What if the RM score rises but quality gets worse? | Reward overoptimization; countermeasures = KL coefficient / early stopping / independent evaluation |
| Why does DPO need no RM and no sampling? | The closed-form optimal policy eliminates the RM, leaving only a classification loss on preference pairs |
| What is the "environment" in RLHF vs. classic RL? | Environment = the language model itself; a single episode; sparse reward (terminal score only) |
Likely follow-ups: interviewers will probably continue with "So how do GRPO and PPO differ in RLHF?", "How do you do online RLHF?", and "Will RLAIF replace human labeling?" The answers are all on this page and the frontier advances page — the key is being able to articulate "environment–reward–policy–exploration" in the LLM context. For the full question bank, see Interview Questions.
Further Reading
- RLHF and Human-Feedback Alignment — detailed mechanics of the three stages, Bradley-Terry, and DPO, plus the RLHF-vs-classic-RL comparison.
- Reward engineering — reward as specification, reward-hacking detection, and preference learning.
- Policy gradient methods — the full mechanics of PPO clipping and GAE (the engine under RLHF).
- Close Reading of Classic Papers — a section-by-section close reading of the InstructGPT paper, with "how to answer this in an interview."
- Frontier advances — the latest evolution of the DPO family, RLAIF, RLVR, and DeepSeek-R1.
- Agents and dialogue systems — RL fine-tuning when an LLM acts as an agent and "reward = task success."
References
- Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback (InstructGPT). arXiv:2203.02155.
- Christiano, P., et al. (2017). Deep Reinforcement Learning from Human Preferences. NeurIPS 2017. (arXiv:1706.03741)
- Stiennon, N., et al. (2020). Learning to Summarize from Human Feedback. NeurIPS 2020. (arXiv:2009.01325)
- Touvron, H., et al. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288. (RLHF details of LLaMA-2-chat)
- Rafailov, R., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290.
- Tunstall, L., et al. (2023). Zephyr: Direct Distillation of LM Alignment. arXiv:2310.16944.
- Lee, H., et al. (2023). RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. arXiv:2309.00267.
- Bai, Y., et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073.
- Schulman, J., et al. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347. (the original PPO paper; RLHF's optimizer)