Skip to content

Alignment: RLHF and DPO

At a glance Alignment is the post-training stage that makes models "both smart and reliable." This article breaks down InstructGPT's SFT→Reward Model→PPO three steps, the KL penalty mechanism, explains why DPO doesn't need a reward model, and compares the trade-offs of RLHF/DPO/RLAIF/Constitutional AI and other methods, concluding with alignment tax and weak-to-strong generalization.

Alignment: RLHF and DPO ​

Alignment is the process of adjusting a language model's behavior through human feedback or preference signals so it aligns with human values, intentions, and safety requirements. Pretraining and SFT answer "can the model do it"; alignment answers "should the model do it" — it turns a "predict the next token" statistical model into an assistant that is "helpful, honest, and harmless."

The classic alignment objective: HHH

The industry's commonly used alignment objective framework is helpful, honest, harmless. These three often pull in opposite directions: too safe seems rigid (hurts helpfulness); too accommodating leads to fabrication (hurts honesty); real alignment engineering is balancing among these three objectives. These three words are also the starting point of Safety & Risks when discussing safety alignment.

Alignment's full pipeline builds on top of fine-tuning — it typically starts with SFT, then layers on preference optimization. This article unfolds chronologically: first the pioneering RLHF (Reinforcement Learning from Human Feedback), then the reinforcement-learning-free DPO (Direct Preference Optimization), followed by new methods and open questions.

1. From SFT to RLHF: InstructGPT's Three Steps ​

The 2022 OpenAI InstructGPT paper (Training language models to follow instructions with human feedback) was the first to systematically demonstrate that "human feedback" can transform a GPT-3 model that "can only continue writing" into one that "can follow instructions," with its 175B model significantly defeating the original GPT-3 in human evaluation. ChatGPT is the productization of this technology — see ChatGPT and Chat Models.

RLHF has three steps: SFT → Train reward model → PPO reinforcement learning.

1. Step One: SFT, Let the Model First Learn to Imitate ​

Use human-written high-quality instruction-response pairs (~13K demonstrations) to supervised fine-tune the pre-trained model. This step teaches the model the format and basic behavior of "answering instructions," but the model still outputs harmful content and fabricates facts — it's just "imitating well."

2. Step Two: Train the Reward Model (RM) ​

The reward model is a "scorer" responsible for ranking and scoring "different responses to the same instruction," turning human preference into an optimizable scalar signal.

Training data: annotators compare multiple responses to the same instruction pairwise ("which is better"), forming preference pairs (win/lose). The reward model is trained on these preference pairs:

text
Input: instruction x and candidate response y
Output: scalar score r(x, y)

Training objective (Bradley-Terry model):
  Make the score of "human-chosen response y_w" higher than "rejected response y_l":

    L_RM = - E_{(x, y_w, y_l) ~ D} [ log σ( r(x, y_w) - r(x, y_l) ) ]

  Intuition: the larger the score difference that aligns with human ranking, the smaller the loss.

The reward model typically initializes from the SFT model, replacing the final output layer with a scalar regression head. Its quality directly determines RLHF's ceiling — if the RM scores well, the policy model learns correctly; if the RM has holes, PPO will exploit them (reward hacking).

3. Step Three: PPO Reinforcement Learning ​

Using the reward model as environmental feedback, update the policy model π_θ via Proximal Policy Optimization (PPO) so its responses "score higher in the RM's eyes":

text
Objective function:
  max  E_{y ~ π_θ(·|x)} [ r(x, y) ]
       └ want the model to output responses that get high reward from RM

  But add two constraints to prevent runaway:

  ① KL penalty: high reward ≠ good response
     max  E [ r(x, y) ] - β · KL( π_θ(y|x) ‖ π_SFT(y|x) )
                           └ don't deviate too far from the SFT model
     Hyperparameter β controls "obedience level": larger β → more conservative, more like SFT model

  ② Mix in raw pretraining data to prevent "only pleasing the RM":
     + γ · E_{x ~ pretrain distribution} [ log π_θ(x) ]

PPO is essentially an online policy: each update round samples responses from the current model → sends to RM for scoring → uses reward for gradient updates. This also means RLHF training is very expensive — it repeatedly runs inference sampling and policy updates, with engineering complexity far exceeding SFT.

KL penalty is the soul of alignment

Without KL penalty, PPO quickly "overfits the reward model": the model learns to output RM-favored boilerplate, stack keywords, or even syntax-broken but extremely high-scoring sentences. KL penalty "tethers the policy model near the SFT model," allowing only limited deviation in the human preference direction. This is RLHF's first gate against reward hacking.

InstructGPT Three Steps Summary ​

StepData/signalModelRole
SFT13K human demonstrationsPolicy model π_SFTLearn basic instruction following
RM33K human preference comparisonsReward model r(x,y)Turn preferences into scores
PPORM feedback + KL penaltyPolicy model π_RLOptimize behavior in the preference direction

4. RLHF's Four Engineering Challenges ​

RLHF has a high effect ceiling, but engineering-wise it's extremely grueling. Four main challenges:

ChallengeManifestationCountermeasure
Reward hackingModel exploits RM loopholes, outputs "high score but bad" responsesKL penalty, periodic RM replacement, red team detection
Reward model overfittingRM saturates on training preferences, loses discriminative powerRM data diversification, continuously add hard samples
PPO training instabilitySampling randomness causes policy fluctuation, training collapseControl KL coefficient, gradient clipping, small step updates
Online sampling costEach update round requires strategy sampling + RM scoringHigh compute budget, long training cycles

This is why "RLHF is expensive and hard" is industry consensus: many teams do daily alignment with DPO, reserving PPO for scenarios that truly require online exploration.

2. DPO: Preference Optimization Without Reinforcement Learning ​

RLHF works well, but engineering-wise it's heavy: train an RM, write a PPO loop, tune KL coefficients, handle stability issues between sampling and updates. In 2023, Rafailov et al.'s DPO paper (Direct Preference Optimization: Your Language Model is Secretly a Reward Model) provided a "surprisingly elegant" alternative.

1. Key Insight: The Reward Model Is "Hidden" in the Optimal Policy ​

DPO's critical observation is mathematical: the RLHF KL-constrained optimization problem has a closed-form solution — the optimal policy π* and the implicit reward function r* have a one-to-one correspondence:

text
r*(x, y) = β · log( π*(y|x) / π_ref(y|x) ) + constant
              └ the log-ratio of the optimal policy is equivalent to a reward function

In other words: "learning a reward function" and "learning a preference-aligned policy" are two sides of the same coin. Since they're equivalent, no need to explicitly train an RM — just feed preference data directly to the policy model for "implicit reward" maximization.

2. DPO Loss ​

DPO writes preference data (x, y_w, y_l) directly into the loss function, letting the policy model π_θ have higher "implicit reward" for the "preferred response y_w" than the "rejected response y_l":

text
L_DPO(π_θ; π_ref) = - E_{(x, y_w, y_l) ~ D} [ log σ( β · log( π_θ(y_w|x) / π_ref(y_w|x) )
                                                     - β · log( π_θ(y_l|x) / π_ref(y_l|x) ) ) ]

Where:
  π_ref is the frozen reference policy (usually the SFT model)
  β is the temperature coefficient, controlling "trust level" of preferences
  σ is the sigmoid function

Intuitively: if the current model increases its probability (relative to π_ref) for y_w more than for y_l, the loss is smaller. The whole training is standard cross-entropy gradient descent — no RM, no PPO, no online sampling.

3. Why DPO Works: Three Keys ​

  1. Implicit reward: the policy's own probability ratio encodes preference information; no need to explicitly model RM.
  2. Offline training: preference pairs are a fixed dataset, stable training, reproducible, memory and compute requirements far below PPO.
  3. No deviation from reference policy: the loss function naturally penalizes "deviating too far from π_ref"; the KL constraint is absorbed into the objective, eliminating the complexity beyond manual β tuning.

DPO's practical status

DPO has become one of the de facto standard alignment methods for the open-source community (e.g., some alignment reports for Llama, Qwen, Mistral series) because it's simple, stable, and low-barrier. But understand its boundaries: it assumes preference data is "statically trustworthy" and can't discover new situations through online sampling like PPO can; in complex scenarios needing "exploratory improvement" (like reasoning capability enhancement), you still need to return to RL.

4. DPO Practical Notes ​

DPO is simple but not no-tuning: preference pair quality directly determines the ceiling (clean first if labeling noise is high); temperature coefficient β affects alignment intensity, too large causes excessive deviation from the reference policy; reference model π_ref is usually the SFT model, consistency needs to be maintained in subsequent iterations. Engineering handling of these details is in Fine-Tuning Practice: Full LoRA Pipeline.

5. RLHF vs DPO Comparison ​

DimensionRLHF (PPO route)DPO
Needs reward model?Yes, train RM separatelyNo, policy is implicit reward
Training methodOnline sampling + RLOffline batch gradient descent
Engineering complexityHigh (PPO loop, KL tuning, stability)Low (standard cross-entropy training)
Memory/ComputeHigh (extra RM + policy sampling)Low
ReproducibilityPoor (sampling randomness)Good (deterministic data)
Exploration capabilityYes (online sampling)No (fixed preference pairs)
Applicable stagePursuing ceiling, data evolves during trainingQuick alignment, open-source community baseline

3. More Alignment Methods: RLAIF, Constitutional AI, KTO, ORPO ​

1. Constitutional AI and RLAIF: Let AI Evaluate AI ​

RLHF's bottleneck is human labeling is expensive and unstable (different annotators have different standards; harmful samples require human review). Anthropic's Constitutional AI (2022) proposed: input a set of manually authored principle list (constitution) into the model, let the model itself "critique-revise-critique again" on its responses, generating harmless preference pairs for preference optimization — humans write only principles, not label each item. The version using AI feedback instead of human feedback is called RLAIF (RL from AI Feedback).

text
Constitutional AI self-revision loop (one example):
  Model's original response A0 (possibly harmful)
  ① Critique: point out A0's problems according to principle "don't provide information that harms others"
  ② Revise: rewrite based on critique, getting A1
  ③ Repeat: A1 critiqued and revised again → A2 (usually now harmless)
  Add (original A0, revised A2) as "bad/good" preference pairs to training data

RLAIF's experimental conclusion: AI feedback achieves effects comparable to human feedback on most metrics (Anthropic 2023), making alignment scalable — this is also the common starting point for later "model self-alignment" methods.

2. KTO: Just "good/bad," No Pairwise Comparison ​

RLHF and DPO both rely on pairwise preferences (which is better). KTO (2024) observed: in real products, what we more easily get is single-item feedback — "this response is good" or "this response is bad" (likes/dislikes, approved/rejected). KTO introduces prospect theory (from Nobel laureates Kahneman & Tversky) into alignment, directly optimizing with binary good/bad signals:

text
KTO intuition:
  - For "good responses": increase their probability (but avoid over-rewarding, aligning with prospect theory's "diminishing returns for gains")
  - For "bad responses": penalize their probability
  Don't need to know "what it's better than," just need "it itself is good/bad"

3. ORPO: SFT and Preference Optimization in One Shot ​

ORPO (2024)'s slogan is "monolithic preference optimization without reference model": it merges "imitating good responses" and "disliking bad responses" into the same loss, eliminating the need for a reference model π_ref, collapsing the training pipeline to a single ordinary fine-tuning pass. The cost: theoretically the stabilizing anchor of a reference distribution is missing; effects compare unevenly with DPO.

Alignment Methods Panorama ​

MethodYearSignal sourceNeeds RM?Needs RL?One-liner
RLHF (PPO)2022Human pairwise preferenceYesYesClassic three-stage, high ceiling, high cost
Constitutional AI / RLAIF2022/2023AI feedback + principlesNo (RM training can be skipped)OptionalReplace manual labeling with principles
DPO2023Human/AI pairwise preferenceNoNoPolicy as reward, simple and stable
KTO2024Single good/bad signalNoNoFits like/dislike real data
ORPO2024Pairwise preferenceNoNoSaves even the reference model

4. Open Questions in Alignment ​

1. Alignment Tax ​

Alignment training (RLHF/DPO and safety fine-tuning) sometimes makes the model worse at certain capabilities — e.g., over-refusal leading to decreased tool usage rate, safety constraints making the model "too cautious" on edge cases. This "capability sacrificed for alignment" is called the alignment tax. It's not inevitable: good alignment engineering (like scaling preference via RLAIF, finely tuning KL coefficients) can keep the tax very low, but the "safety vs capability trade-off" tension persists long-term. This is also why Evaluation & Benchmarks insists on "looking at both capability and safety metrics before/after alignment."

2. Weak-to-Strong Generalization ​

OpenAI's 2023 paper poses an academic question facing "super-alignment": when future models become strong enough that humans can't reliably judge their output quality, can we still align them? Experiments found: when using "weak model (human-level)" labels to supervise "strong model," the strong model often performs better than the weak supervisor itself (weak-to-strong generalization) — but it doesn't fully inherit all signals from the weak supervisor. This hints that "using AI to assist supervising AI" may be the only feasible route, providing a theoretical footnote for RLAIF's long-term viability.

3. Alignment Evaluation Itself Is Hard ​

Alignment effects ultimately depend on evaluation. Unlike capability evaluation, alignment evaluation targets "whether behavior matches expectations," using more composite means:

Eval MethodWhat it tests
Human blind A/B win rateProportion of "better" responses after alignment (InstructGPT's core metric)
HHH graded scoringScore by helpful / honest / harmless dimensions
Safety red teamJailbreak success rate, violation rate (see Safety & Risks)
Preference consistencyHow well model behavior aligns with annotator preferences
Refusal rate monitoringCorrect refusal rate, false refusal rate (measure of alignment tax)

Alignment isn't a one-time training step; it's a continuous loop of "train-evaluate-redteam-retrain"; specific engineering methods are in Evaluation in Practice. Remember one principle: alignment without evaluation is just self-satisfaction.

Further Reading ​

References ​