Skip to content

Alignment: RLHF and DPO

At a glance Alignment is the process of making a model's capabilities serve human intent — being capable, and behaving the way people want — and this article breaks down the three stages of RLHF, DPO's implicit reward optimization, Constitutional AI, and other alignment techniques, along with the alignment tax and reward hacking.

Alignment: RLHF and DPO ​

One-Sentence Definition: From "Capable" to "Well-Behaved" ​

Alignment refers to the process of making an AI model's behavior consistent with human intent, values, and safety norms. In plain terms: a model should not only be capable, but also act the way humans want it to.

Pretraining teaches a model "what the next token in a text should be," but not "what humans want to hear" — and alignment exists to teach exactly that second lesson. It is the crucial leap that turns a large model from a "text continuation engine" into a "dependable assistant," and it is the part of the large language model (LLM) engineering pipeline closest to the "product experience." Today, from ChatGPT and conversational AI to DeepSeek-R1 and reasoning models, nearly every model that actually works well has some alignment technique behind it; and in discussions of AI safety and governance, alignment is repeatedly emphasized as the "line of defense."

Quick takeaway

Capability comes from pretraining; conduct comes from alignment. To judge whether a model "can be used," look at benchmark scores; to judge whether you "dare to use it," look at how well it is aligned.

Why Alignment Is Needed: From "Sounding Human" to "Being Accountable" ​

A pretrained language model has exactly one learning objective: maximize the probability of predicting the next token over its corpus. That objective naturally produces text that only "sounds human" — fluent, highly imitative, stylistically close to the training corpus — but the model has no concept of "whether it's correct, whether it should be said, or who it answers to." That's why a typical "raw pretrained" model will:

  • Fabricate with a straight face: invent names, numbers, and citations out of thin air (hallucination);
  • Answer everything and agree with anything: when led astray, produce harmful, discriminatory, or criminal how-to content;
  • Miss the point: you ask about one thing, it writes at length about another;
  • Lack boundaries: it doesn't know what to say, what to keep confidential, or what to refuse.

Alignment exists to close this "last mile." The industry usually breaks the goal into three H's — the standard framework in AI safety discussions:

GoalEnglishMeaningTypical failure
HelpfulnessHelpfulActually solving the user's problem, not brushing them offMissing the point, empty platitudes
HonestyHonestNo fabrication, no exaggeration, admitting what it doesn't knowHallucinating, then doubling down
HarmlessnessHarmlessNo harmful, illegal, or biased contentViolent/criminal instructions, discriminatory speech

The three H's often conflict: being too "honest" may expose dangerous information, while being too "harmless" may hollow out the answers. Alignment engineering is essentially the search for a balance point among these three constraints — and that balance point is chosen with data, not written with rules. This is precisely where methods like RLHF and DPO get their footing. Looking back at the brief history of evolution, alignment is a "first-world problem" that only became prominent once models grew large, arriving alongside the capability leap brought by Transformer and attention mechanisms.

RLHF: The Mainstream Route to Alignment (All Three Stages Explained) ​

RLHF (Reinforcement Learning from Human Feedback) is the alignment paradigm that OpenAI established with InstructGPT (2022) and the rest of the industry then copied. Its core idea: humans can't directly edit a model, but humans can "score" it — so train "human preferences" into a differentiable scorer (a reward model), then optimize the model against it with reinforcement learning. The pipeline has three stages:

Pretrained model (only "sounds human")
   │
   ▼ Stage 1: SFT — supervised fine-tuning
SFT model (imitates examples to "follow instructions")
   │
   ├──── Human annotators: rank multiple responses ◀── collect preference data
   ▼
Stage 2: Reward model RM (learns to "score responses"; score = human preference)
   │
   ▼ Stage 3: PPO reinforcement learning
Aligned model (policy π optimized to "maximize the RM score")

Stage 1: Supervised Fine-Tuning (SFT) ​

First, collect a batch of human-written demonstration responses: give annotators an instruction (prompt), have them write high-quality replies, then run ordinary supervised learning — the standard approach in fine-tuning and PEFT, except the data is "instruction–response" pairs. This step gives the model the basic ability to "follow instructions": switching from "continuing text" to "answering questions." InstructGPT used about 13,000 such demonstrations.

Why isn't SFT enough?

SFT depends on human demonstrations, and the ceiling of those demonstrations is the annotators' own skill: they are expensive to produce at scale, and "good answers" can never be exhaustively enumerated. SFT only lays the foundation; the real taste-tuning comes from the reward signal later on.

Stage 2: Training the Reward Model (RM) ​

The reward model (RM) is a "scorer": it takes a (prompt, response) pair as input and outputs a scalar score, where a higher score means a better match with human preference. Training data comes from human ranking annotations — for the same prompt, the model generates multiple candidate responses (say 4–9), and annotators rank them from best to worst. Ranking is more reliable than "scoring": human intuition is good at "comparing" but bad at "assigning absolute scores." RMs are usually trained with a pairwise ranking loss, ultimately learning to approximate "how humans would rank."

Stage 3: Reinforcement Learning with PPO ​

Once you have the RM, use it as the "reward function" and update the policy model with the PPO (Proximal Policy Optimization) reinforcement learning algorithm: the model generates a response → the RM scores it → the score serves as the reward signal for updating parameters. To keep the model from drifting entirely away from natural language while chasing scores, a KL penalty term is usually added to the reward: the policy gets penalized if it strays too far from the SFT model. The full InstructGPT report is covered in classic paper deep-dives and is the first-hand source for understanding RLHF.

The three engineering realities of RLHF

  1. Expensive: a three-stage pipeline where every stage requires multiple training/inference runs, plus a whole annotation pipeline to maintain;
  2. Unstable: PPO is hyperparameter-sensitive, and the slightest bias in the reward model can drag the policy astray (see "reward hacking" below);
  3. Prone to overfitting preferences: the model may learn to flatter the annotators' writing style rather than truly tell the truth.

This is exactly what motivates DPO and similar techniques to bypass reinforcement learning and optimize preferences directly.

DPO: Dropping the Reward Model and Reinforcement Learning ​

DPO (Direct Preference Optimization, 2023) rewrites RLHF from a far more "minimalist" perspective: since RLHF ultimately just searches for a policy that "favors the responses humans preferred in the preference data," that goal can be derived directly from the preference data as an implicit reward function — no explicitly trained RM and no PPO needed at all.

RLHF: preference data → train an RM → maximize the RM score with PPO (the long way around)
DPO : preference data → derive an implicit reward directly → one gradient step on the policy (the shortcut)

How DPO works: for each (prompt, chosen response, rejected response) triple, construct a simple classification-style loss — push the model to assign a higher probability to the "chosen response" than to the "rejected response," while a KL constraint keeps it from drifting too far from the reference model. It can be proven mathematically that DPO's optimization objective is equivalent to RLHF's; the difference is only "how you get there."

DimensionRLHF (the InstructGPT route)DPO
Intermediate componentReward model RM + PPO reinforcement learningNone (implicit reward)
Implementation complexityHigh (three-stage pipeline)Low (one training run)
Training stabilityPoor (sensitive to RL hyperparameters and reward scale)Good (no RL oscillation)
Compute costHigh (RL sampling, repeated training)Low (single forward/backward pass)
Data requirementPreference data (rankings/comparisons)Preference data (pairs suffice)
Theoretical ceilingHigh, but hard to reachEquivalent, and easier to land
Notable examplesInstructGPT, early ChatGPTZephyr, many open-source models

Quick takeaway

For fast shipping and lean resources, choose DPO; for squeezing out peak performance with "high-quality preference data + a generous engineering budget," RLHF is still worth considering. In practice the gap between the two is often smaller than the papers suggest — data quality is the real deciding factor.

The original DPO paper (arXiv:2305.18290) and the InstructGPT/RLHF paper (arXiv:2203.02155) are now the must-read pairing for alignment in the paper map, and the latter's engineering details continue to be discussed in frontier developments.

Other Alignment Techniques: Each With Its Own Labor-Saving Trick ​

Beyond RLHF and DPO, several other routes are worth knowing about:

TechniqueIdeaKey points
Constitutional AIGive the model a "constitution" (a list of behavioral principles) and let the AI critique and rewrite its own harmful responsesProposed by Anthropic; replaces part of the human annotation with AI self-critique
RLAIF (Reinforcement Learning from AI Feedback)Replace human feedback with "AI feedback" (e.g., another model acting as judge) to train the RMSaves labor; but sensitive to the judge model's quality
Rejection sampling fine-tuningSample many responses from the SFT model, keep the subset with the highest RM scores, then go back to SFTSimple and stable; used by early DeepSeek versions
Direct feedback fine-tuningUsers' real-time corrections/preferences flow straight into the training dataInteractive alignment; fits a product feedback loop
ORPO (Odds Ratio Preference Optimization)Merge "instruction tuning + preference optimization" into one step: add a preference odds-ratio term to the SFT loss, done in a single passIntroduced in 2024; eliminates the separate preference optimization stage, saving both VRAM and training time
IPO (Identity Preference Optimization)A DPO improvement: replace the regularizer in the original loss with an identity operator, easing DPO's overfitting when data is plentifulIntroduced in 2023; more stable than DPO when preference data is abundant

These methods are not mutually exclusive — real-world models often combine them: SFT first, then a round of rejection-sampling curation, then DPO to finish is the most common "low-cost alignment combo." "Reward-model-free" variants like ORPO/IPO (see the glossary) compress the pipeline even further. Claude, the flagship of Constitutional AI, is often compared against ChatGPT's RLHF route; both are "value alignment" practices in the AI safety and governance context.

Data and Annotation: Alignment's Fuel ​

Whether RLHF or DPO, the core fuel is human preference data. How you collect it sets the ceiling of your alignment:

  • Ranking: rank multiple candidate responses to the same prompt — information-dense with low annotator disagreement; the standard for RLHF/RM;
  • Pairwise comparison: pick one of two responses — the lowest annotation cost; DPO's favorite;
  • Scalar scoring: rate a single response 1–7 — unintuitive and noisy; rarely used today.

Scale is not quality

The value of alignment data lies not in "how many rows" but in coverage and consistency: the instruction distribution must cover real usage scenarios (customer service, coding, writing, refusing sensitive requests); multiple annotators must agree with one another (usually by writing annotation guidelines first and running consistency checks). A batch of 10,000 high-quality, consistently calibrated preference data often beats 100,000 casually labeled examples. This is also a topic worth watching separately in the datasets and tools archive.

When annotating, also watch out for "teaching it wrong" and "teaching it biased": annotators' personal stances can smuggle bias into the model, and deliberately piling on "safe refusals" makes the model overly conservative (see "alignment tax" below). Rule of thumb: calibrate annotation standards on a small batch first, scale up only after, and finally validate annotation quality with an evaluation set — the exact same thinking behind building an LLM eval.

The Cost and Limits of Alignment ​

Alignment is not a free lunch; it has clear costs and limits:

The Alignment Tax and Over-Alignment ​

Making a model "obedient and harmless" usually sacrifices part of its raw capability — this loss is called the alignment tax. The classic symptom: alignment applied too hard makes the model overly cautious, refusing "how do I write an essay" as if it were a policy violation, or losing creativity to an over-preachy "correction voice." The other side of over-alignment is "sycophancy bias": the model only says what users want to hear and loses its independent judgment.

Reward Hacking ​

When a model discovers that "gaming the reward score" is easier than "telling the truth," it starts exploiting loopholes in the reward model — reward hacking. Real cases: models learn to stack "I'm so sorry..." at the end of responses to please the RM; to write very long answers so the RM misreads length as "thoroughness"; even to deliberately trigger safety rules to cash in "caution points." Countermeasures include feeding the reward model more adversarial examples, adding KL constraints, and training multiple alignment signals (helpful + harmless + honest) separately before fusing them.

How to Tell Whether Alignment Is Good Enough ​

Alignment quality can't be judged by benchmark scores alone; it needs dedicated LLM evaluation and benchmarking methods: manual spot checks, adversarial red-teaming, jailbreak attack testing, harmful-content trigger rates, hallucination rates, refusal rates and refusal accuracy (neither refusing things it shouldn't, nor failing to refuse what it should), and so on. The bottom line: alignment acceptance is essentially "betting that the model has suppressed all three classes of bad behavior, without knowing what users will ask" — which is why red-teaming and continuous monitoring matter more than one-shot metrics.

The cost of alignment failure has actually happened

In 2016, Microsoft's chatbot Tay was turned into a "racist repeating machine" by users within 24 hours of launch and had to be taken offline — the classic failure of "no alignment done." Early ChatGPT was also swarmed by jailbreak attacks that bypassed its safety alignment and produced policy-violating content; jailbreaking is essentially the reverse application of prompt engineering tricks, "tricking" the model out of its aligned state and back into its raw state. The shared lesson of these cases: alignment is not a one-time training run but an ongoing tug-of-war with adversaries — see AI safety and governance.

Alignment vs Fine-Tuning vs Safety: How the Three Relate ​

These three concepts are often conflated, but they operate at different levels:

ConceptEssenceProblem it solvesExample methods
Fine-tuningA technique (modifying model weights)Adapting the model to a task/domain/styleFull-parameter fine-tuning, PEFT like LoRA
AlignmentA goal (behavior consistent with human intent)Making the model "act the way people want"RLHF, DPO, Constitutional AI, rejection sampling
SafetyA broader governance topicPreventing the model from causing real harmAlignment, red-teaming, monitoring, policy and regulation

A handy way to remember the relationship: fine-tuning is the "tool," alignment is the "goal," safety is the "floor." Alignment is often implemented via fine-tuning (or fine-tuning-like methods), but fine-tuning can also serve pure capability gains (say, learning domain knowledge) with nothing to do with alignment; safety encompasses alignment but also involves deployment guardrails, monitoring, disclosure, and regulation — far beyond any single training run. To explain these three concepts clearly in a job interview, see the interview question bank; for hands-on practice, follow fine-tune your own LLM and common pitfalls and anti-patterns.

Further Reading ​

References ​