Skip to content

LLM Alignment: RLHF in Practice

On this page The three-stage training of InstructGPT/ChatGPT: SFT -> reward model -> PPO; open-source reproductions (LLaMA-2-chat, Zephyr); the real costs of reward overoptimization and the alignment tax; how DPO simplifies it all; large-model training from an RL engineer's perspective.

LLM Alignment: RLHF in Practice ​

In a nutshell: this page explains, from an RL engineer's perspective, what RLHF (reinforcement learning from human feedback) actually is — how the three-stage SFT → reward model → PPO pipeline works, why the reference model and the KL penalty are necessary, the true costs of reward overoptimization and the "alignment tax," and how DPO compresses all of this into a single idea. It is the full worked example accompanying the RLHF and Human-Feedback Alignment concept page.

1. Background: Pretrained Models "Continue Text" — They Don't "Converse" ​

1.1 Where the Problem Comes From ​

By 2020, GPT-3 could "continue" any text — but it had three cardinal sins:

SymptomHow it shows upRoot cause
Won't follow instructionsAsk it to "help me write an email" and it may just continue the sentence instead of actually answeringThe pretraining objective is "next-token probability," not "get the task done"
Harmful / biased outputIt produces discriminatory, offensive, or false contentIt learned the bad parts of internet corpora
Confidently wrongIt fabricates facts and delivers wrong answers with a straight faceThere is no "right/wrong" signal — only "does it sound like human text"

The fundamental tension: pretraining optimizes for "more like the corpus," while a product needs "more helpful, more truthful, safer" — two different objectives. The latter is what the field calls alignment.

1.2 Why "Just Train Another SFT" Isn't Enough ​

Supervised fine-tuning (SFT) does help — fine-tune on human-written "instruction–answer" pairs and the model learns what an answer looks like. But SFT has two ceilings:

  • Data volume: high-quality human-labeled instruction–answer pairs are expensive, capping out in the tens of thousands — nowhere near covering the behavior space;
  • The imitation ceiling: SFT only imitates the annotators' "demonstrations" — it never beats them and can't handle "edges the demonstrations didn't cover." This is the "limits of behavior cloning" from the multi-armed bandits page, magnified to the scale of large models.

So in the 2022 InstructGPT paper, OpenAI delivered the full answer: turn "human preference" into an optimizable reward signal, then optimize it with reinforcement learning (PPO) — that is RLHF.

2. The Three-Stage Pipeline: SFT → Reward Model → PPO ​

text
Stage 1  SFT: supervised fine-tune      Stage 2  RM: reward model          Stage 3  RL: PPO optimization
┌──────────────────────────────┐     ┌──────────────────────────────┐     ┌──────────────────────────────┐
│ Humans write                 │     │ Model generates              │     │ RM score as the reward;      │
│ instruction-answer pairs;    │     │ multiple answers; humans     │     │ PPO updates the policy;      │
│ fine-tune w/ cross-entropy   │──►  │ rank/pair them; train        │──►  │ KL penalty against the       │
│ (learn the answer shape)     │     │ a scorer RM (predicts        │     │ frozen reference model       │
│ tens of thousands of samples │     │ human preference)            │     │ (learn better answers)       │
└──────────────────────────────┘     └──────────────────────────────┘     └──────────────────────────────┘
StageInputTraining objectiveResources required
SFTHuman instruction–answer pairsMaximize the log probability of the answersTens of thousands of labeled samples
RM (reward model)Multiple answers to the same prompt + human preferencesGive "more preferred" answers higher scoresHundreds of thousands of preference pairs
PPOPrompt pool + RM scoresMaximize "reward − KL penalty"Substantial inference compute

TIP

In RLHF, the "environment" is the language model itself: the action is "generate a token / an answer," the state is "the sequence of tokens generated so far," and the reward is the RM's final score. This maps one-to-one onto the "RLHF vs. classic RL" table on the RLHF and Human-Feedback Alignment concept page — it is an episodic, one-shot sequential decision problem with no intermediate rewards, only a terminal score.

3. Training the Reward Model: The Bradley-Terry Intuition ​

3.1 Why Humans Can't Be the Reward Directly ​

If a human had to judge "is this right?" after every generated token, they'd burn out — and their labels would be inconsistent anyway. RLHF's solution is to train a surrogate reward function (RM): have humans compare answers in batches, then learn a scoring model from those comparisons.

3.2 Preference Data and Bradley-Terry ​

Given the same prompt x, the model generates two answers — y_w (the better one) and y_l (the worse one) — and a human labels which is better. The output of the reward model r_θ(x, y) is interpreted as the probability that a human prefers a given answer, via the Bradley-Terry model:

text
P(y_w preferred over y_l) = σ( r_θ(x, y_w) − r_θ(x, y_l) )

σ = sigmoid function
Loss: −log σ( r_θ(x, y_w) − r_θ(x, y_l) )
Intuition: push the score of "the answer humans like" well above "the answer they don't"

The RM is typically a sequence model that emits a scalar score at the final token. InstructGPT trained it on roughly 33,000 prompts, each with multiple answers, yielding preference data at that scale.

3.3 Key Details of Preference Data ​

  • Within-prompt comparisons: comparisons must be between "two answers to the same prompt"; cross-prompt comparisons are meaningless;
  • Labeler consistency: measure agreement among annotators on the same answer pair (e.g., Cohen's κ);
  • Ranking vs. scoring: "ranking" is a more robust annotation format than "absolute scores" — humans are better at ordering things than scoring them.

This whole recipe of "learning a reward from pairwise preferences" is exactly the deployed form of the "preference learning: deriving rewards from human preferences" section on the reward engineering page.

3.4 Reward Model Generalization and Calibration ​

The RM is what PPO directly optimizes, so its quality sets the ceiling for RLHF. Three engineering points:

PointPracticePitfall
Cover the training distributionPreference prompts must cover "the prompts the policy will see"Train only on safe prompts and the RM scores dangerous ones wildly
Reward mean calibrationRe-score after each new model generation; track RM score driftSystematically inflated RM scores → misjudged overoptimization
Stay in sync with the policyWhen the policy changes, re-validate the RM's "preference judgments"An old RM evaluating a new policy produces distorted assessments

A pragmatic check: during PPO training, periodically measure the RM's accuracy on a batch of human-labeled answers — if its accuracy on that data drops below 70%, something is wrong with the preference data or the RM itself; fix the RM before training the policy further. This is exactly the "independent evaluation set" idea from the evaluation and benchmarks page.

4. The PPO Stage: Four Models on Stage at Once — Actor / Critic / Reference / RM ​

4.1 The Roles of the Four Models ​

PPO fine-tuning is the heaviest engineering step in RLHF — four models run at the same time:

text
actor     π_θ        : the model being trained; produces answers (the policy)
reference π_ref      : the frozen SFT model; measures "how far the policy has drifted" (KL baseline)
reward    r_φ        : the RM trained in stage 2; scores answers
critic    V_ψ        : value network; estimates "expected reward−KL from the current state" (needed by PPO)

────────────────────────────────────────────────────────────
Training loop (per prompt):
1. The actor samples an answer (a complete episode)
2. The RM scores the whole answer: r_φ(x, y)
3. Per-token KL penalty: r̂_t = r_φ(x,y) only at the final token − β·KL(π_θ ‖ π_ref)
4. Update the actor with PPO; update the critic with GAE

4.2 Why the Reference Model and KL Penalty Are Needed ​

Maximize the RM score directly, and the model will find the RM's loopholes (reward hacking) — saying whatever fools the RM into awarding high points. The reference model π_ref (usually the SFT model) provides an "anchor of reasonable language," penalizing the policy per token by its KL distance from that anchor:

text
KL penalty term: −β·D_KL( π_θ(·|s_t) ‖ π_ref(·|s_t) )

Meaning: every time the policy "drifts from the reference model's distribution," points are deducted.
Effect: the model is allowed to improve, but not to lose the ability to "talk normally" in exchange for reward.

The intuition: this is the answer to the "why KL?" section on the RLHF concept page — the KL penalty is RLHF's "exploration guardrail": it prevents the policy from degenerating entirely (babbling nonsense to farm reward) and from distribution collapse (only ever saying one thing). Seen from the exploration vs. exploitation angle, the KL constraint is RLHF's distinctive form of exploration: exploring within a neighborhood of the reference distribution.

4.3 Hyperparameters and Training Realities ​

ItemTypical values (InstructGPT / open-source reproductions)Notes
KL coefficient βOn the order of 0.01–0.1Larger = more conservative
PPO clip range0.2Same as classic PPO
Learning rateSeparate for actor/critic, on the order of 1e-6 to 1e-5Fine-tuning large models requires baby steps
BatchHundreds to thousands of promptsOne answer sampled per prompt
Training stepsA few thousandAlignment doesn't need many steps; overtraining degrades

Alignment training is "lightweight" in scale: the InstructGPT paper showed that 1.3B InstructGPT beat the 175B unaligned GPT-3 in human evaluations — alignment makes "small models more useful," rather than relying on parameter count.

4.4 Implementation Notes for PPO on LLMs ​

PPO in RLHF differs from PPO in robotics or games in three implementation details that are easy to trip over:

Implementation pointWhat it meansCommon pitfall
Reward only at the final tokenThe RM score for the whole answer is added only to the last token; all earlier tokens get zero rewardSpreading the RM score across every token breaks the semantics
Per-token KLThe KL penalty is computed distribution-wise at each token position, not once for the whole answerSkipping per-token KL leads to severe overoptimization
Reference model stays frozenπ_ref is never updated; it is the "anchor"Accidentally training π_ref too disables the KL constraint
Critic mirrors the actorThe value network usually copies the actor architecture (shared backbone or standalone)A critic that is too small estimates GAE poorly

WARNING

One more point that even RL veterans tend to miss: RLHF "episodes" are extremely short (an answer of a few dozen to a few hundred tokens), so the semantics of GAE's λ differ from those in long-horizon environments. In practice, many open-source reproductions skip fine-tuning λ entirely — the default works well enough. But if you're doing RLHF on long-answer tasks (e.g., long-document summarization), revisit λ and reward normalization.

5. Case Studies, Data, and Results ​

5.1 Human Evaluation of InstructGPT ​

MetricResult
1.3B InstructGPT vs. 175B GPT-3Humans preferred InstructGPT's outputs (by a significant majority)
HelpfulnessFar ahead of the unaligned model
TruthfulnessHallucinations reduced, though not eliminated
ToxicityMitigated on adversarial prompts

5.2 LLaMA-2-chat: A Complete Open-Source Loop ​

Meta released LLaMA-2 in 2023, and LLaMA-2-chat disclosed complete RLHF details:

  • Multi-round RLHF: SFT → two or more "reward model + PPO" iterations; in each round, preference labels are re-collected on samples from the previous round's model, the RM is updated, and the policy is retrained;
  • A family of reward models: separate "helpfulness RM" and "safety RM" trained on different prompt distributions;
  • Rejection sampling fine-tuning: filter sampled answers through the RM, then SFT directly on the high-reward samples before PPO — an "offline distillation" trick that reduces PPO's exploration cost;
  • The alignment tax, documented: Meta reported that after RLHF, capability stayed flat or dipped slightly on most tasks while safety and helpfulness improved — public evidence for the "alignment tax."

5.3 Zephyr: Simplified Open-Source Alignment with DPO ​

HuggingFace's Zephyr-7B (Tunstall et al., 2023) proved the engineering value of DPO (Direct Preference Optimization): no RM training, no PPO — just one "closed-form" fine-tune on preference pairs — and a 7B model matched or beat stronger open-source models of the day on benchmarks like MT-Bench.

5.4 Dataset Composition and the Alignment Cost Ledger ​

RLHF's true cost is routinely underestimated; break it down to see:

StageMain costRule-of-thumb share
Preference data labelingHumans labeling hundreds of thousands of preference pairsHigh (labor)
RM trainingModerate computeLow
PPO sampling and trainingMultiple samples + four model forward passesHigh (compute)
Evaluation and regressionIndependent evaluation set + human spot checksMedium

InstructGPT's approach was to sample prompts from real API user requests (rather than hand-writing prompts), covering the true distribution; each prompt got multiple answers, which annotators compared pairwise, producing a dataset of ~33,000 prompts and hundreds of thousands of preference pairs. This lesson is worth copying: for preference data, "distribution coverage" matters more than "raw volume" — sampling from real usage is far more efficient than inventing prompts in a vacuum.

WARNING

Alignment cost can't be judged by a single training run: as models iterate, preference data goes stale (it labels outputs of older models) and must be continuously re-sampled and re-labeled. The cost of maintaining a "preference data pipeline" over the long run is often several times that of any single training run. This is the hidden bill every alignment team runs into.

6. The Alignment Tax and Reward Overoptimization: Goodhart's Law ​

6.1 Two Real Costs ​

Alignment tax: alignment training can make the model worse at tasks it used to be good at (math, code) — because preference data favors "likable" over "correct." LLaMA-2's report of slightly lower scores on some benchmarks after RLHF is exactly this cost.

Reward overoptimization: as PPO steps pile up, the RM score keeps climbing while true human evaluation scores start falling — the model is "farming the RM's points" instead of getting better. This is Goodhart's law ("when a measure becomes a target, it ceases to be a good measure") in its standard RLHF form:

text
             RM score
               ↗↗↗
  training steps ────►
               ↘↘
             true human-preference score (rises, then falls)

Countermeasures:

  1. KL penalty: smaller β is not better — too small and you get overoptimization;
  2. Early stopping: decide when to stop using human evaluation or a held-out RM validation set — don't keep training just because the RM score is still rising;
  3. DPO's "implicit KL": DPO bakes the KL constraint into the objective itself, making it inherently less prone to overoptimization than PPO (though not immune).

WARNING

The most common RLHF failure in industry is "the RM score is still going up, so let's train a few more rounds." The RM score is a proxy metric, not the real objective; the real objective is "humans find it useful." The right approach is an evaluation set + human spot checks independent of the training RM, treating "RM score plateauing or dropping" together with "human evaluation dropping" as the early-stopping signal. This corresponds exactly to the "don't make decisions with a single metric" discipline on the evaluation and benchmarks page.

7. DPO: Compressing RLHF into One Sentence ​

7.1 The Core Idea of DPO ​

The key insight of Rafailov et al. (2023): the optimal solution of "reward maximization + KL constraint" in PPO can be written in closed form in terms of the reward function r(x, y). This lets you eliminate PPO and the RM, turning "preference" directly into a classification loss:

text
Optimal policy for PPO's objective (maximize reward − β·KL):
  π*(y|x) ∝ π_ref(y|x) · exp( r(x,y) / β )

Solve for the reward in reverse:
  r(x,y) = β·log[ π_θ(y|x) / π_ref(y|x) ] + constant

Substitute r back into the Bradley-Terry preference loss -> the DPO loss:
  L = −log σ( β·log[ π_θ(y_w|x)/π_ref(y_w|x) ] − β·log[ π_θ(y_l|x)/π_ref(y_l|x) ] )

Intuition: DPO no longer trains a reward model explicitly; it directly requires the policy to raise the relative probability of y_w and lower that of y_l (relative to the reference model). It keeps RLHF's "KL constraint + preference objective" while eliminating all the engineering complexity of RM training and PPO.

7.2 PPO vs. DPO ​

DimensionPPO-RLHFDPO
Reward modelMust be trained separatelyNot needed
Training stabilitySensitive (four models, many hyperparameters)Simple and stable
KL constraintExplicit β coefficientImplicit (reference model in the objective)
Sample efficiencyRequires large amounts of online samplingOffline preference data suffices
ScalabilityStrong at online feedback / iterative alignmentExcels offline; online variants emerging
Open-source ecosystemLLaMA-2-style RLHF pipelinesWidely adopted by Zephyr, Mistral-family models

7.3 Which One, When ​

  • Want "good chat experience, simple engineering": DPO is enough — Zephyr is the proof;
  • Need "online iteration" (the model generates, preference data is continuously refreshed): online RLHF in the PPO family is the better fit;
  • Research frontier: the two are converging (online variants of DPO, GRPO, etc.) — see Frontier Advances.

7.4 The Full Variant Landscape: PPO / DPO / IPO / KTO / GRPO ​

After 2023, alignment algorithms exploded — one table to see the mainstream variants:

AlgorithmCore ideaNeeds RM?Needs reference model?Characteristics
PPO-RLHFRM scoring + KL penalty + policy optimizationYesYesMost general; heavy engineering
DPOClosed-form solution eliminates the RM and samplingNoYesSimple, stable, offline
IPOFixes DPO's overfitting (uses preference pairs directly)NoYesSteadier, simpler sampling
KTOOnly needs "good/bad" labels, no pairingNoYesLower data bar
GRPORelative rewards within a group, drops the critic (adopted by DeepSeek)NoYesSaves memory; suits reasoning RL

Quick selection guide: paired preference data + simplicity → DPO; only good/bad labels → KTO; large-scale reasoning RL (math/code) → the GRPO route; online iteration → the PPO family.

8. The Frontier: RLAIF, Online RLHF, and Reasoning Alignment ​

8.1 RLAIF: Replacing Human Feedback with AI Feedback ​

RLAIF (RL from AI feedback): instead of having humans label preferences, a strong LLM (say, GPT-4) acts as the "judge," generating preference labels that feed the same training pipeline. Early work from DeepMind and the public study of Lee et al. (2023) showed that AI feedback can approach human-feedback quality on tasks like harmlessness, at far better cost and scalability than human labeling. It turns RLHF from "labor-intensive" into "scalable."

8.2 Online RLHF and Reasoning Alignment (RLVR) ​

The hottest direction in 2024–2025 is using RL directly to improve LLM reasoning:

  • RLVR (RL with verifiable rewards): for tasks where answers can be verified automatically (math, code), use a rule-based verifier (correct/incorrect) as the reward, bypassing the RM entirely — the reward is objective, with no dependence on human preference;
  • DeepSeek-R1 and friends: large-scale RL teaches the model "chain-of-thought self-reflection," yielding huge gains in math/code reasoning — with no SFT cold-start stage (pure RL straight from the base model);
  • This direction expands RLHF from "aligning human preferences" to "optimizing verifiable capabilities," and it is the star of the "RL for reasoning" section on the frontier advances page.

8.3 Convergence with Agent RL ​

When an LLM is used as an agent (calling tools, browsing the web, executing multi-step tasks), the "reward" shifts from "human preference" to "did the task succeed" (was the tool call correct? did the web action achieve the goal?) — the RLHF framework carries over unchanged, with only the reward source swapped. That's the subject of the Agents and Dialogue Systems page.

9. Key Takeaways for RL Engineers ​

9.1 RLHF vs. Classic RL ​

Classic RLRLHFCorrespondence
EnvironmentThe language model itself (no real environment)Environment = policy
RewardEnvironment signalRM score (proxy reward)
Policyπ(a|s)Language model π(y|x)
Explorationε-greedy, entropyKL to the reference model (exploration guardrail)
Sparse rewardsCommonRLHF has a single terminal score (sparsest possible)
Value estimationQ/V networksCritic network
Engineering focusEnvironment and dataPreference data and training stability

9.2 What RL Engineers Can Bring Over ​

  • All the PPO machinery transfers directly: GAE, clipping, value clipping, gradient clipping;
  • Evaluation discipline: multiple seeds, independent evaluation sets, early stopping — all just as valid in RLHF;
  • Reward-engineering instincts: the RM is "the reward function," and the detection methods for reward hacking (see reward engineering) are called "reward overoptimization" in the LLM world — same thing.

9.3 Common Pitfalls (From an RL Perspective) ​

PitfallSymptomCountermeasure
KL coefficient tuned haphazardlyOutputs become perfunctory or unhingedRun ablations on β; don't settle for the default
RM out of sync with the actorInflated RM scoresRecalibrate the RM regularly on fresh model outputs
OvertrainingHuman evaluations dropIndependent evaluation set + early stopping
Dirty preference dataThe RM learns the wrong thingLabel-consistency checks, prompt diversity
Treating the proxy metric as the goalUsers unhappy after launchClose the loop with human evaluation / live metrics

9.4 Five Must-Know RLHF Interview Questions (with Answer Points) ​

RLHF is a high-frequency topic in RL job interviews. Here are the five inevitable questions, with answer points:

QuestionAnswer points
Why does RLHF need three stages — why not PPO directly?You need a reward function first (the RM); the RM needs preference data; PPO with no signal goes nowhere
Who gets the KL penalty, and why?Applied between the policy and the reference model, to prevent reward hacking and distribution collapse
What if the RM score rises but quality gets worse?Reward overoptimization; countermeasures = KL coefficient / early stopping / independent evaluation
Why does DPO need no RM and no sampling?The closed-form optimal policy eliminates the RM, leaving only a classification loss on preference pairs
What is the "environment" in RLHF vs. classic RL?Environment = the language model itself; a single episode; sparse reward (terminal score only)

Likely follow-ups: interviewers will probably continue with "So how do GRPO and PPO differ in RLHF?", "How do you do online RLHF?", and "Will RLAIF replace human labeling?" The answers are all on this page and the frontier advances page — the key is being able to articulate "environment–reward–policy–exploration" in the LLM context. For the full question bank, see Interview Questions.

Further Reading ​

References ​

  • Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback (InstructGPT). arXiv:2203.02155.
  • Christiano, P., et al. (2017). Deep Reinforcement Learning from Human Preferences. NeurIPS 2017. (arXiv:1706.03741)
  • Stiennon, N., et al. (2020). Learning to Summarize from Human Feedback. NeurIPS 2020. (arXiv:2009.01325)
  • Touvron, H., et al. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288. (RLHF details of LLaMA-2-chat)
  • Rafailov, R., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290.
  • Tunstall, L., et al. (2023). Zephyr: Direct Distillation of LM Alignment. arXiv:2310.16944.
  • Lee, H., et al. (2023). RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. arXiv:2309.00267.
  • Bai, Y., et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073.
  • Schulman, J., et al. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347. (the original PPO paper; RLHF's optimizer)