Skip to content

RLHF and Alignment with Human Feedback

On this page The key technique that makes language models "speak human" — the limits of behavior cloning, reward model training, and the three-stage PPO pipeline; the alignment tax, reward overoptimization, DPO and the simplification of direct preference optimization; how RLHF compares with classic RL.

RLHF and Alignment with Human Feedback ​

In a nutshell: this page is all about RLHF (reinforcement learning from human feedback) — the key technique that turns a large language model from "able to continue a sentence" into "able to talk like a human": the three-stage SFT → reward model → PPO pipeline, why the KL constraint is indispensable, how reward overoptimization happens, and how DPO collapses it all into a single step. By the end you'll be able to draw the full RLHF data flow, explain the role of every component, and articulate how it compares with classic RL.

1. The alignment problem: why "pretraining + SFT" is not enough ​

1.1 The problem: capability is not conversation ​

A pretrained model's job is to "predict the next token" — it knows how to continue a piece of text, but not what a good reply actually looks like. Supervised fine-tuning (SFT) teaches it to "follow instructions" using human-written (instruction, desired response) pairs.

But SFT has three hard limitations:

Limitation of SFTExplanation
Imitation ceilingIt can only imitate the responses in the SFT data — it never rises above the quality of the examples
Data is expensiveHigh-quality, human-written instruction–response pairs are extremely costly
Cannot express "preference"The data only says "this response is good", never "of these two responses, which is better, and by how much"

The last point matters most: real human preferences are relative, multi-dimensional, and hard to capture in a single example. Ask for "a paper abstract" and two AI responses may both be decent while one is clearly better. SFT has no way to encode this kind of "relatively better".

1.2 The core insight of RLHF ​

Instead of giving the model the "correct answer", give it a "comparison signal": which of two responses is better. Comparison data is far cheaper to produce than written demonstrations, and it encodes human preferences natively.

RLHF (the InstructGPT/ChatGPT recipe, Ouyang et al., 2022) makes alignment a matter of "optimizing with RL a reward learned from human preferences". This is why it earns a place in an RL handbook: RLHF is the canonical RL instance where the reward is defined by human preference — the environment is the token-generation process, and the reward is the reward model's score.

2. The three-stage pipeline: SFT → reward model → PPO ​

text
RLHF: three stages (the InstructGPT recipe)
┌────────────────────────────────────────────────────────────┐
│ Stage 1: SFT (supervised fine-tuning)                      │
│   data: (instruction, human-written ideal response)        │
│   output: π_SFT — a model that "follows instructions"      │
│                                                            │
│ Stage 2: Reward model (RM)                                 │
│   data: (instruction, response A, response B,              │
│          human label "which is better")                    │
│   output: r_φ(x, y) — a model that scores responses        │
│                                                            │
│ Stage 3: PPO fine-tuning                                   │
│   models: policy π_θ (initialized from π_SFT)              │
│           value net V_ψ + reward model r_φ + ref π_ref     │
│   objective: max E[r_φ] − β·KL(π_θ ‖ π_ref)                │
│   output: the aligned model π_θ                            │
└────────────────────────────────────────────────────────────┘

2.1 Stage 1: SFT (the starting point of behavior cloning) ​

Supervised fine-tuning on a few thousand to tens of thousands of human-written (instruction, high-quality response) pairs. The purpose of this stage is to "pull the model back onto the rails of instruction following"; it is the behavioral foundation for everything that follows. The quality of the SFT data directly sets the ceiling — and, as covered in Policy Gradient Methods, it is behavior cloning applied at scale, with the limitations discussed above.

2.2 Stage 2: the reward model (turning preferences into scores) ​

The reward model $r_\phi(x, y)$ takes (instruction x, response y) and outputs a scalar score. Its training data consists of preference pairs:

instruction x: "Explain quantum entanglement in one sentence"
response A: "Quantum entanglement is..."   ← the human labeler picks A as better
response B: "Entanglement is entanglement"

The Bradley–Terry model writes the "preference probability" as a sigmoid of the reward difference:

$$ P(A \succ B \mid x) = \frac{\exp(r_\phi(x, A))}{\exp(r_\phi(x, A)) + \exp(r_\phi(x, B))} = \sigma\big(r_\phi(x,A) - r_\phi(x,B)\big) $$

Training objective: maximize the log-likelihood of the human-labeled preference pairs:

$$ \mathcal{L}{RM} = -\mathbb{E}{(x, A, B) \sim D}\left[ \log \sigma\big(r_\phi(x,A) - r_\phi(x,B)\big) \right] $$

The intuition: a reward model is a "learned scorer" — it distills human "which is better" judgments into a differentiable function that scores any response. That function then becomes the reward source for RL optimization.

Why preference pairs instead of direct scores

Asking people to score responses from 1 to 5 gives inconsistent scales (one annotator's 4 is another's 3) and is mentally taxing. Binary choices are cheaper and more consistent — relative judgments are far more reliable than absolute ones. This is the same idea as the "preference learning" section in Reward Engineering.

2.3 Stage 3: PPO fine-tuning (RL inside "prompt sampling") ​

This step ports classic PPO onto a language model:

  • Environment = a prompt x, plus the model's token-generation process;
  • Action = the generated response y (a sequence of tokens);
  • Reward = the reward model's score $r_\phi(x,y)$ (sparse: one score per response);
  • Policy = the LLM itself, $\pi_\theta$.

The optimization objective (the core formula of RLHF):

$$ \max_\theta \underbrace{\mathbb{E}{x \sim \mathcal{D}, y \sim \pi\theta}[r_\phi(x,y)]}{\text{make the reward high}} - \underbrace{\beta , D\big( \pi_\theta(y|x) ,|, \pi_{ref}(y|x) \big)}_{\text{don't drift too far}} $$

Here $\pi_{ref}$ is the reference model (usually π_SFT or the previous version of the policy), and β is the KL-penalty coefficient.

2.4 Why the reference model and KL penalty are mandatory ​

In one sentence: the reward model is only "locally trustworthy" — it has judgment only near "normal responses". Once the policy generates text far from the training distribution (gibberish, broken syntax, extreme wording), the reward model's scores are completely unreliable (an out-of-distribution problem; see Offline RL). The KL penalty does the following:

Role of KLExplanation
Caps the exploration radiusThe policy is not allowed to drift too far from the reference model (see Section 8 of Exploration and Exploitation)
Prevents reward-model hallucinationStops the policy from optimizing into regions where the reward model mistakenly gives high scores
Preserves language qualityPrevents collapse into keyword-stuffing "reward farming"
Provides an entropy-like constraintSimilar to entropy regularization in classic RL, keeping outputs diverse

What happens if you remove the KL term

In training you will soon see "reward soaring, language collapsing" — the policy discovers a hole in the reward model (for example, endlessly repeating "quantum entanglement quantum entanglement quantum entanglement..." to harvest points). The KL penalty is the non-negotiable seatbelt of RLHF engineering. Remove it and you get the reward hacking from Reward Engineering all over again — this time on a massive model.

3. The cast of models on stage in RLHF training ​

During PPO, four models are actually running at once — the heaviest part of RLHF engineering:

ModelRoleNeeds gradients?
Policy model π_θ (actor)Generates responses; gets optimizedYes
Value model V_ψ (critic)Estimates the state value of the current generation sequence (for the advantage)Yes
Reward model r_φScores responsesNo (frozen, inference only)
Reference model π_refThe reference for computing the KL penaltyNo (frozen, inference only)

The training loop (the PPO-ptx variant from InstructGPT):

text
loop:
  1. Sample a prompt x ~ D
  2. Generate a response y with π_θ (one rollout)
  3. reward = r_φ(x,y) + KL penalty (against π_ref)
  4. Compute the advantage (GAE, needs V_ψ)
  5. Update π_θ (clipped objective + entropy bonus)
  6. Update V_ψ (TD regression)
  7. Optional: mix in 10% SFT data to prevent "catastrophic forgetting"
     (the PPO-ptx trick)

RLHF through the eyes of an RL engineer

The PPO stage of RLHF is a standard on-policy actor-critic; the only special feature is that the reward comes from a learned reward model, and each episode is a single response — one sparse reward. GAE is nearly vacuous here (there are no intermediate-step rewards), and the advantage is provided mostly by the "step-by-step accumulation of the KL penalty". This is why PPO implementations for RLHF are often trimmed down to something quite lightweight.

4. The alignment tax and reward overoptimization (Goodhart's law) ​

4.1 The alignment tax ​

Alignment tax: the alignment process (RLHF) costs the model points on standard capability benchmarks (such as MMLU or code evaluations). "Pleasing human preferences" and "general capability" do not fully align — over-catering makes the model "more polite but more mediocre". The InstructGPT paper reports this capability regression explicitly and mitigates it by "mixing in SFT data".

4.2 Reward overoptimization: higher reward ≠ a better model ​

This is the most critical empirical phenomenon in RLHF: the longer training runs, the higher the reward-model score climbs — but human ratings of the actual output rise first and then fall.

text
  quality ▲
          │        ╱‾‾‾‾╲
          │     ╱/       ╲        ← human rating (true quality)
          │   ╱/
          │  ╱
          ├──────╲──────────────► training steps
          │        ╲  reward model score (keeps climbing)
          │         ╲
          └──────────────────────────────►
              true quality starts declining past a point,
              while the reward model still "thinks" things are improving.

This is Goodhart's law: "When a measure becomes a target, it ceases to be a good measure." The reward model is a "proxy for human preference"; optimizing it to the extreme means optimizing a proxy that has drifted away from the real objective.

4.3 Engineering countermeasures ​

TechniqueMechanism
KL constraintCap the exploration radius; don't let the policy leave the reward model's trustworthy region
Early stoppingMonitor the reward-vs-true-quality divergence curve and stop before the turning point
Reward model ensembles / calibrationHave multiple RMs vote; calibrate with human spot checks
Mixed objectivesPPO-ptx mixes in SFT data to prevent forgetting
Preference-label spot checksPeriodically evaluate the final policy with humans instead of trusting RM scores alone

5. DPO: treating preferences as the objective directly (bypassing the reward model) ​

5.1 Motivation: RLHF is heavy engineering ​

RLHF requires maintaining four models plus a PPO training loop; the engineering is complex and training is unstable (the reward model, KL, and GAE all need tuning). DPO (Direct Preference Optimization, Rafailov et al., 2023) proved that the reward-model stage can be eliminated mathematically.

5.2 Core idea: a closed-form translation of "preferences" into a policy objective ​

The key insight of DPO: the RLHF optimization problem (maximize reward − KL penalty) has a closed-form optimal solution:

$$ \pi^*(y|x) \propto \pi_{ref}(y|x) , \exp\left( \frac{1}{\beta} r_\phi(x,y) \right) $$

Invert this to recover the reward: $r(x,y) = \beta \log\frac{\pi^*(y|x)}{\pi_{ref}(y|x)} + \text{const}$. Substituting "reward = the policy-to-reference ratio" back into the Bradley–Terry preference loss makes the reward model disappear entirely, yielding the DPO loss:

$$ \mathcal{L}{DPO}(\theta) = -\mathbb{E}{(x, A, B) \sim D}\left[ \log \sigma\left( \beta \log \frac{\pi_\theta(A|x)}{\pi_{ref}(A|x)} - \beta \log \frac{\pi_\theta(B|x)}{\pi_{ref}(B|x)} \right) \right] $$

The intuition: DPO trains the policy directly on preference pairs — make the policy more likely to produce the "preferred A" and less likely to produce the "rejected B", with the KL ratio controlling the size of the change (β is the KL strength). No reward model, no PPO.

5.3 RLHF vs. DPO ​

DimensionRLHF (PPO route)DPO
ComponentsSFT + RM + PPO (4 models)SFT + preference pairs (2 models)
Training stabilityPoor (PPO is finicky)Good (simple classification-style training)
HyperparametersMany (β, KL, clip…)Few (β)
ScalabilityHeavy, expensiveLight, cheap
Robustness to reward off-distributionKL as the safety netAlso constrained by the β ratio
Online samplingYes (RL-style)Mostly offline (on-policy sampling variants exist)
Representative modelsInstructGPT, ChatGPT, Llama 2-ChatZephyr and the DPO family

How to choose today

  • With engineering resources, chasing the ceiling: RLHF (especially online RLHF — see below); better results and easy to iterate;
  • Fast alignment, open-source replication: the DPO family (Zephyr et al.) — up and running within a day;
  • The academic and industry consensus: the DPO simplification is the direction of travel, but whether you need "PPO-grade RL" depends on your budget and quality bar. Frontier variants include IPO, KTO, ORPO, and more (see Frontier Advances).

6. RLHF vs. classic RL ​

DimensionClassic RLRLHF
EnvironmentGrids / robots / gamesThe token-generation process
State sEnvironment statePrompt + tokens generated so far
Action aJoint torques / directionsThe next token
Transition PPhysics / game rulesThe model's own autoregressive sampling
Reward RHand-designed (see Reward Engineering)Scores from a learned reward model
SparsityOften sparseExtremely sparse (one score per response)
Explorationε / entropy / intrinsic rewardsKL constraint + sampling temperature
Data sourceOnline trial and errorHuman preference labels + self-sampling
EvaluationReturn curvesHuman evaluation + RM scores + capability benchmarks

A curious inversion: in classic RL we hand-write the reward (because the task is quantifiable); in RLHF we learn the reward (because preferences resist being put into words). The two converge on the unified framework of "optimizing return" — RLHF is a direct instance of the classic RL framework applied to language generation.

7. Limitations and frontiers ​

7.1 Inherent limitations of RLHF ​

LimitationExplanation
Preference-label biasAnnotators are subjective; position and order biases exist
Reward overoptimizationGoodhart's law; needs early stopping + KL (see Section 4)
Aligns only to "average human preference"Diverse users and minority needs get averaged away
Optimizes for "looking good"No guarantee of factual correctness (hallucination still needs external fixes such as retrieval augmentation)
High engineering cost4 models + PPO + distributed RL training

7.2 Frontier directions ​

  • RLAIF (RL from AI feedback): use another LLM as the "labeler" to generate preference pairs in place of human annotators, greatly improving scalability (Bai et al., 2022, Constitutional AI);
  • Online RLHF: alternate reward-model and policy training (e.g., self-play-style preference sampling), mitigating the staleness of offline preference data;
  • RLHF for reasoning (RLVR): replace preferences with "verifiable rewards" (does the code pass the tests, is the math answer right) — when the rules can verify, no human feedback is needed and you do straight RL (the DeepSeek-R1 route); see Frontier Advances;
  • The whole DPO variant family: IPO, KTO, SimPO, ORPO, and more, each trading off simplicity against robustness.

7.3 Historical lineage ​

RLHF did not appear out of thin air: it inherits the actor-critic of Value-Based Learning, the PPO of Policy Gradient Methods, and the preference learning of Reward Engineering. For the full history and the classic papers, see A Brief History of RL and Classic Papers, Close Reading (the Ouyang InstructGPT entry).

8. Key takeaways for engineers ​

  1. SFT is the foundation: neither the reward model nor PPO can rescue a model trained on poor SFT data;
  2. Reward-model quality sets the ceiling: effort spent on preference-data cleaning and labeling consistency pays off far more than tuning PPO hyperparameters;
  3. The KL penalty is non-negotiable and β must be tuned: β too large → you revert to SFT (no learning); β too small → reward hacking;
  4. Monitor the divergence curve: reward vs. human spot-check quality, and stop before the turning point;
  5. DPO first, RLHF second: on a limited budget, get a baseline with DPO first, then invest in RLHF to push the ceiling.

For practical details (open-source replication, three-model training, cost accounting), see LLM Alignment in Practice: RLHF; for the conceptual discussion of reward design and hacking, see Reward Engineering.

Further Reading ​

References ​

  • Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training language models to follow instructions with human feedback. NeurIPS. arXiv:2203.02155 (InstructGPT — the standard citation for RLHF)
  • Rafailov, R., Sharma, A., Mitchell, E., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS. arXiv:2305.18290 (DPO)
  • Bai, Y., Jones, A., Ndousse, K., et al. (2022). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv:2204.05862 (Constitutional AI / RLAIF)
  • Stiennon, N., Ouyang, L., Wu, J., et al. (2020). Learning to summarize with human feedback. NeurIPS. arXiv:2009.01325 (the precursor work to RLHF)
  • Ziegler, D. M., Stiennon, N., Wu, J., et al. (2019). Fine-Tuning Language Models from Human Preferences. arXiv:1909.08593 (an early version of RLHF)
  • Christiano, P. F., Leike, J., Brown, T. B., et al. (2017). Deep Reinforcement Learning from Human Preferences. NeurIPS. arXiv:1706.03741 (where the RLHF concept was introduced)