Modern AI Engineering

Lesson 8.10 · 25 min

DPO: Alignment Without the Separate Reward Model

What if the whole reward-model-plus-PPO machinery of RLHF could be replaced by one loss function you train like ordinary fine-tuning?

In short: DPO (Direct Preference Optimization) trains a language model straight from preference pairs (a prompt, a chosen answer and a rejected answer) with a simple classification-style loss. It uses a frozen reference model and a math result showing that the RLHF objective’s optimal policy defines an implicit reward, so no separate reward model and no reinforcement learning loop are needed. It is cheaper and more stable than PPO-based RLHF, but it learns only from fixed data and does not explore.

What is RLHF and why do we need it?

A model after supervised fine-tuning (SFT) can answer questions, but we want it to give the answers people prefer: helpful, honest, safe and in the right tone. RLHF (reinforcement learning from human feedback) does this in two extra steps: train a reward model on human comparisons (“answer A is better than answer B”), then use PPO, a reinforcement learning algorithm, to make the model produce answers the reward model scores highly, with a KL penalty that keeps it close to the SFT model.

Running example: our bike-rental support assistant. We have collected thousands of comparisons where a support lead picked the better of two replies to a customer message. We want the model to learn from them.

The problem with RLHF

  • Many moving parts: a policy, a reference model, a reward model and a value model must live in memory at the same time.
  • Sampling during training: PPO must generate fresh text in every iteration, which is slow.
  • Instability: results depend on many hyperparameters (learning rates, clip range, KL coefficient, GAE settings) and small implementation details.
  • Reward hacking: the policy can exploit flaws in the separately trained reward model.

Many teams with good preference data simply could not afford or stabilise this pipeline. That was the motivation for DPO.

What is Direct Preference Optimization?

DPO was introduced by Rafael Rafailov, Archit Sharma, Eric Mitchell and colleagues at Stanford in 2023, in a paper titled “Direct Preference Optimization: Your Language Model is Secretly a Reward Model”. It optimises the same goal as KL-regularised RLHF but does it directly on preference pairs, using one supervised-style loss. No reward model is trained, and no text is sampled during training.

Think of it like learning from marked exam pairs RLHF is like training a separate examiner, then having the student write new essays for the examiner to grade, over and over. DPO skips the examiner: the student looks at pairs of essays where a teacher already marked which was better and adjusts directly, making the better one feel more “like me” and the worse one less, compared with how they wrote before the course.

What is preference data?

A preference dataset is a list of triples (x, y_w, y_l): a prompt x, a chosen (winning) response y_w, and a rejected (losing) response y_l. The pair is usually two responses to the same prompt, judged by humans or by a strong AI model.

One preference example for our support bot
FieldContent
Prompt x“My e-bike battery died halfway through my rental. What now?”
Chosen y_w“Sorry about that! Use the app’s Help → Battery issue to get a free swap at the nearest station, and we’ll credit the lost time.”
Rejected y_l“Batteries can die for many reasons, including temperature, age and usage patterns. Lithium-ion cells…” (long and does not help)

This is exactly the kind of data RLHF’s reward model is trained on, so the same datasets work for both. Public examples include Anthropic’s HH-RLHF and UltraFeedback.

The key idea behind DPO

RLHF wants the policy that maximises reward minus a KL penalty. That problem has a known exact solution: π*(y | x) = π_ref(y | x) · exp(r(x, y) / β) / Z(x), where Z(x) is a normalising constant. In words, the best policy reweights the reference model toward higher-reward answers.

DPO’s insight is to run this equation backwards. Rearranging gives the reward in terms of the policy:

Now plug this into the Bradley–Terry preference model, which only looks at the difference in reward between two answers to the same prompt. The awkward β · log Z(x) term is identical for both answers, so it cancels. What remains depends only on the policy and the reference, both of which we can compute. So we can train the policy directly on preferences: the language model is, in the paper’s phrase, secretly a reward model.

Pause and think: Why is it essential that the Z(x) term cancels?

Z(x) sums over every possible response, which is impossible to compute for a language model. Because both answers in a pair share the same prompt x, Z(x) appears in both rewards and disappears in their difference, leaving only quantities we can compute: log-probabilities under the policy and the reference.

The DPO loss function in simple words

Read it in three parts:

  • Implicit reward of each answer: β times how much more (or less) likely the policy makes the answer than the reference does.
  • Margin: chosen reward minus rejected reward. We want this positive and large.
  • Loss: −log σ(margin), the same logistic loss used to train RLHF reward models. It is about 0.69 at margin 0 and falls toward 0 as the margin grows.

The gradient has a helpful built-in weighting: each pair’s update is scaled by σ(−margin), so the model works hardest on pairs it currently ranks wrongly and barely touches pairs it already gets right. The reference model matters too: without it, the easiest way to lower the loss would be to shift probabilities wildly; the ratio against π_ref plays the role of the KL penalty.

dpo_loss_steps.py

import numpy as np
beta = 0.1
# Log-probabilities (summed over tokens) of the chosen and rejected answers
ref_chosen, ref_rejected = -42.0, -40.0     # frozen reference (SFT) model
pol_chosen, pol_rejected = -42.0, -40.0     # policy starts as a copy of ref
def sigmoid(x): return 1 / (1 + np.exp(-x))
for step in range(4):
# implicit rewards: how much more the policy likes each answer than ref does
r_c = beta * (pol_chosen - ref_chosen)
r_r = beta * (pol_rejected - ref_rejected)
margin = r_c - r_r
loss = -np.log(sigmoid(margin))
weight = sigmoid(-margin)               # big when the model is still wrong
print(f"step {step}: logp chosen {pol_chosen:6.1f}  rejected {pol_rejected:6.1f}"
f"  margin {margin:+.2f}  loss {loss:.3f}  grad weight {weight:.2f}")
# gradient of the loss pushes chosen up and rejected down (toy step of size 20)
pol_chosen += 20 * beta * weight
pol_rejected -= 20 * beta * weight

Output:

step 0: logp chosen  -42.0  rejected  -40.0  margin +0.00  loss 0.693  grad weight 0.50
step 1: logp chosen  -41.0  rejected  -41.0  margin +0.20  loss 0.598  grad weight 0.45
step 2: logp chosen  -40.1  rejected  -41.9  margin +0.38  loss 0.521  grad weight 0.41
step 3: logp chosen  -39.3  rejected  -42.7  margin +0.54  loss 0.458  grad weight 0.37

Pause and think: At step 0 the reference model already likes the rejected answer more (−40 vs −42), yet the margin is 0. Why?

DPO’s margin uses log-ratios to the reference, not raw log-probabilities. The policy equals the reference at step 0, so both log-ratios are 0. DPO measures how the policy has changed relative to the reference, not absolute likelihoods.

How DPO works step by step

A DPO training run

  1. Start from an SFT model: Fine-tune on good demonstrations first. This model becomes both the starting policy and the frozen reference.
  2. Collect preference pairs: Prompt, chosen, rejected. Ideally the responses come from the SFT model itself, so the data matches the policy.
  3. Precompute reference log-probs: Run the frozen reference once over every chosen and rejected answer and store log π_ref. This can be cached.
  4. Compute policy log-probs: For each batch, run the policy on chosen and rejected answers (a normal forward pass, no sampling).
  5. Apply the DPO loss and update: Compute the margin, the loss −log σ(β·margin), backpropagate, and step the optimiser. Usually 1–3 epochs.
  6. Evaluate: Check win rates against the SFT model with humans or a judge model, plus safety and general-skill tests.

DPO vs RLHF (PPO)

In the original paper, DPO matched or exceeded PPO-based RLHF on the tasks tested (sentiment control, summarisation and single-turn dialogue). Later studies found that well-tuned online methods can still beat DPO in some settings, so the comparison depends on data, scale and tuning.

Advantages and disadvantages of DPO

  • Simple: a few lines of loss code on top of normal fine-tuning.
  • Cheap: no reward model to train, no value model, no generation during training.
  • Stable: behaves like supervised learning, with mainly β and the learning rate to tune.
  • Widely adopted: used in open models such as Zephyr and Tülu 2, and Meta reported using DPO in Llama 3’s post-training.
  • Offline only: it cannot discover better answers than those in the dataset. Iterative or online DPO (regenerate pairs with the current model, relabel, repeat) partly fixes this.
  • Distribution mismatch: if pairs were written by a very different model, the log-ratios can behave poorly.
  • Likelihood can fall for both answers: DPO only cares about the gap, so the chosen answer’s probability can drop too, as long as the rejected one drops more.
  • Overfitting and verbosity: it can overfit small datasets and inherit biases in the data, such as preferring longer answers.

Common mistakes Skipping SFT and running DPO on a base model; using a reference model different from the starting policy; setting β too small so the policy drifts far from the reference; training many epochs on a small dataset; and judging success by training loss instead of win rate on held-out prompts. Variants such as IPO, KTO, ORPO and SimPO each target some of these issues.

Common mistakes and how to spot them

A DPO run can report a falling loss and still produce a worse model. The loss only sees the gap between the chosen and the rejected answer, so we need to know what it hides. One worked case shows the main trap. Numbers are illustrative.

A falling loss with a falling chosen answer

  1. Start: The reference gives the chosen answer a log-probability of −42 and the rejected one −40. The policy is a copy, so the margin is 0 and the loss is ln 2 ≈ 0.693.
  2. After training: The policy now gives the chosen answer −45 and the rejected answer −50. Both went down.
  3. Shifts: Chosen: −45 − (−42) = −3. Rejected: −50 − (−40) = −10.
  4. Margin and loss: With β = 0.1: margin = 0.1 × (−3 − (−10)) = 0.7. Loss = −log σ(0.7) ≈ 0.40. The loss improved from 0.693 to 0.40.
  5. What really happened: The chosen answer is now e⁻³ ≈ 0.05 times as likely as before. The probability that left both answers went to other text, which no pair in the dataset describes.
  6. How to catch it: Log the average chosen log-probability next to the loss. If it keeps sinking while the loss falls, sample some outputs and read them before training further.
Diagnosing a DPO run
What we seeLikely causeWhat to check
Loss stays at 0.693The policy is not moving, or the “reference” is being updated together with the policy so every shift is 0Confirm the reference is a separate frozen copy; confirm the policy has trainable weights
Loss falls close to 0 within a few stepsLearning rate too high, or a tiny dataset being memorisedRead sampled outputs for repetition or nonsense; lower the learning rate
Loss falls, chosen log-probability falls tooBoth answers are being pushed down, as in the example aboveRaise β, lower the learning rate, or stop earlier
High accuracy on training pairs, about 50% on held-out pairsMemorised the pairs instead of learning the preferenceHold out pairs from the start; add data or train for fewer epochs

The useful number here is pair accuracy: the share of pairs with a positive margin. It is easy to explain to a teammate, and on held-out pairs it tells us whether the preference generalises.

Practice: try it yourself

We will score a small batch of four preference pairs the way a DPO training step does. For each pair we compute how far the policy has moved from the reference on each answer, the margin, the loss, and the weight that pair gets in the next update. The four pairs are chosen to show four different situations.

practice_dpo_batch.py

import numpy as np
beta = 0.1
# Four preference pairs. Summed log-probs: (chosen, rejected)
ref = np.array([[-42.0, -40.0],     # frozen reference model
[-30.0, -35.0],
[-55.0, -50.0],
[-20.0, -21.0]])
pol = np.array([[-36.0, -47.0],     # policy after some DPO training
[-30.0, -35.0],     # pair 2: policy has not moved at all
[-58.0, -49.0],     # pair 3: policy moved the WRONG way
[-26.0, -47.0]])    # pair 4: both answers became less likely
shift = pol - ref                              # log(pi / pi_ref) per answer
margin = beta * (shift[:, 0] - shift[:, 1])    # chosen shift minus rejected shift
loss = -np.log(1 / (1 + np.exp(-margin)))      # -log sigma(margin)
weight = 1 / (1 + np.exp(margin))              # sigma(-margin): gradient weight
print("pair  chosen shift  rejected shift  margin   loss  weight")
for i in range(4):
print(f"{i + 1:>4}  {shift[i, 0]:12.1f}  {shift[i, 1]:14.1f}  {margin[i]:6.2f}  {loss[i]:5.3f}  {weight[i]:6.3f}")
print(f"mean loss: {loss.mean():.3f}")
print(f"pairs ranked the right way (margin > 0): {int((margin > 0).sum())} of 4")
print(f"loss of an untrained policy: ln 2 = {np.log(2):.3f}")

Output:

pair  chosen shift  rejected shift  margin   loss  weight
1           6.0            -7.0    1.30  0.241   0.214
2           0.0             0.0    0.00  0.693   0.500
3          -3.0             1.0   -0.40  0.913   0.599
4          -6.0           -26.0    2.00  0.127   0.119
mean loss: 0.494
pairs ranked the right way (margin > 0): 2 of 4
loss of an untrained policy: ln 2 = 0.693

Pair 2 sits exactly at the untrained value 0.693. Pair 3 is worse than untrained. Pair 4 has the lowest loss of all, although the policy made both of its answers less likely.

Now change it:

  • Change beta on line 3 from 0.1 to 0.5. Predict: which pair’s loss rises, which pairs’ losses fall, and what happens to pair 2?
  • Change pair 4 on line 12 to [-26.0, -27.0], so both answers dropped by the same 6. Predict its margin and loss before running.
  • Repair pair 3 on line 11 by setting it to [-52.0, -53.0]. Predict the new “ranked the right way” count and whether the mean loss drops below 0.4.

Pause and think: Pair 4 has the lowest loss (0.127). If we now sample replies for that prompt, is the chosen answer more likely to appear than before training?

No. Its log-probability fell by 6, so it is about e⁻⁶ ≈ 0.0025 times as likely as under the reference. The loss is low only because the rejected answer fell much further (by 26). DPO guarantees nothing about where the freed probability goes. For this prompt we should read actual samples, not trust the loss.

Pause and think: The mean loss (0.494) is clearly better than the untrained 0.693, yet only 2 of 4 pairs are ranked the right way. How can both be true, and which number would we show a teammate?

The mean is pulled down by two confident pairs (0.241 and 0.127), which more than offsets the one bad pair (0.913). Loss rewards being very right on some pairs; accuracy counts each pair once. We would report both, on held-out pairs: accuracy to say how often the preference is respected, and loss to see confidence. A good loss with poor accuracy means the model is over-fitting a subset.

Key takeaways

  • DPO learns from (prompt, chosen, rejected) pairs with one supervised-style loss.
  • It solves the same KL-regularised objective as RLHF: the policy’s log-ratio to the reference is an implicit reward.
  • Loss = −log σ(β·[log-ratio chosen − log-ratio rejected]); wrongly ranked pairs get the biggest updates.
  • Only two models, no sampling, no reward model: cheaper and more stable than PPO.
  • It is offline: no exploration, sensitive to data match, and can lower both likelihoods.

Key terms

  • DPO: Direct Preference Optimization: training a policy directly on preference pairs without RL.
  • Preference pair: A prompt with a chosen (better) and a rejected (worse) response.
  • Reference model: A frozen copy of the starting (SFT) model used to measure how the policy has changed.
  • Implicit reward: β · log(π(y|x) / π_ref(y|x)), the reward a policy implicitly assigns to an answer.
  • Margin: Implicit reward of the chosen answer minus that of the rejected answer.
  • β (beta): Controls how far the policy may move from the reference; larger means more conservative.

← 8.9 PPO: The Reinforcement Algorithm Behind Instruction Tuning · 8.11 GRPO: Group-Based Preference Optimization Explained →