Lesson 8.8 · 27 min
RLHF: Aligning LLMs with Human Preferences
If people can easily tell which of two answers is better but cannot write down a formula for “good answer”, how do we train a model to give better answers?
In short: RLHF (reinforcement learning from human feedback) aligns a language model with human preferences in three stages: supervised fine-tuning on good examples, training a reward model on human comparisons, and reinforcement learning (usually PPO) to maximise that reward. A KL penalty keeps the model close to its starting point so it does not exploit the reward model’s flaws, a failure called reward hacking.
What is RLHF?
RLHF stands for reinforcement learning from human feedback. It is a way to fine-tune a model using human judgements (“this answer is better than that one”) instead of only human-written examples. The judgements train a reward model, a network that scores answers, and reinforcement learning then trains the language model to earn high scores.
In earlier lessons we met where it came from: learning rewards from clip comparisons for robots (2017) and InstructGPT (2022), which applied it to GPT-3. This lesson is the general, practical recipe: each stage, the KL penalty, reward hacking, mistakes and best practices.
Think of it like a cooking show with a trained judge First the cook studies recipes (supervised fine-tuning). Then we train a food critic by showing them pairs of dishes and telling them which one diners preferred (reward model). Finally the cook keeps cooking, the critic scores each dish, and the cook adjusts (RL). A house rule says “do not stray too far from the recipe book” (KL penalty), so the cook cannot win by finding one weird trick the critic overrates.
Why we need RLHF
A pretrained model predicts likely text, and supervised fine-tuning (SFT) teaches it to imitate good answers. Both stop short of what we want:
- “Good” is hard to define in a loss. Helpfulness, tone, honesty and safety do not have a formula. But humans can compare two answers quickly and fairly consistently.
- Imitation has a ceiling. SFT copies demonstrations; it does not learn that one acceptable answer is better than another, and it cannot exceed its demonstrators’ quality.
- SFT sees only good examples. It never learns from the model’s own mistakes. RL lets the model try answers, see which score low, and move away from them.
- Comparisons are cheaper than writing. Ranking four answers takes far less effort than writing one perfect answer.
Running example: our bike-rental support assistant. After SFT it answers politely but sometimes over-apologises, rambles, or promises refunds we do not offer. Humans can easily pick the better of two replies, so RLHF can turn those picks into steady improvement.
The big picture
Four models are involved during Stage 3: the policy being trained, the frozen reference model (the SFT copy), the frozen reward model, and a value model (critic) that PPO uses to estimate expected reward. Holding all four in GPU memory is a big part of why RLHF is expensive.
Stage 1: supervised fine-tuning (SFT)
We collect prompts and high-quality responses (written by people or carefully selected) and fine-tune the pretrained model on them with the normal next-token loss, usually only on response tokens. For our bot: real customer messages paired with ideal replies from our best support agents.
SFT does two jobs. It teaches the basic format (answer the question, use the chat template), and it gives the RL stage a sensible starting point. RL works by improving what the model already sometimes does; if the model never produces a decent answer, there is nothing good for the reward to reinforce.
Stage 2: training the reward model
Building the reward model
- Sample responses: For each prompt, generate two or more responses from the SFT model (sometimes from several models for variety).
- Collect comparisons: Human labelers pick the better response, or rank several. Clear guidelines matter: what counts as helpful, honest and safe.
- Build the model: Start from the SFT model (or a similar LM) and replace the next-token output layer with a head that outputs one number.
- Train with a pairwise loss: For each pair, push the score of the chosen response above the rejected one: loss = −log σ(r(x, y_chosen) − r(x, y_rejected)).
- Validate: Measure accuracy on held-out comparisons. Agreement in the 60–75% range is common for chat data, because humans themselves often disagree.
Small numbers: if the reward model gives the chosen reply 1.4 and the rejected one 0.4, the difference is 1.0 and σ(1.0) ≈ 0.73, so the loss is −log 0.73 ≈ 0.31. If it had the order backwards (0.4 vs 1.4), σ(−1.0) ≈ 0.27 and the loss is about 1.31. Training lowers the loss by widening the right gap.
Pause and think: Only the difference between two scores enters the loss. What does that imply about the reward model’s absolute scores?
They have no fixed meaning: adding the same constant to every score leaves the loss unchanged. That is why reward scores are often normalised (for example to mean 0) before RL, and why a score of “3.2” alone says little.
Stage 3: RL fine-tuning with PPO
Now we treat the language model as a policy π: given a prompt (the state), it generates tokens (actions). After the full response, the reward model gives a score. PPO (Proximal Policy Optimization) then updates the policy so that responses that beat expectations become more likely, while clipping each update to stay small and stable. PPO gets its own lesson next; here we focus on how it fits into RLHF.
- Rollout: sample a batch of prompts and let the current policy generate a response for each.
- Score: the reward model scores each full response; compute the per-token KL penalty against the reference model.
- Advantage: the value model estimates how much reward was expected; the advantage = actual minus expected tells each token whether it did better or worse than usual.
- Update: run a few epochs of PPO’s clipped update on this batch, then repeat with fresh rollouts.
The KL penalty
The reward model is only accurate on the kind of responses it was trained on. If the policy wanders far from them, its scores become unreliable, and RL is very good at finding unreliable high scores. The KL penalty keeps the policy close to the reference (SFT) model:
The log-ratio is computed per token and summed; it is positive when the policy makes a response more likely than the reference did. The best possible policy for this objective has a neat closed form: π(y) ∝ π_ref(y) · exp(r(y) / β). Large β keeps π almost equal to π_ref; small β piles all probability onto the highest-scoring response, flaws and all.
kl_penalty_demo.py
import numpy as np
# Four possible replies to one support question
replies = ["short correct", "detailed correct", "flattering fluff", "wrong but confident"]
pi_ref = np.array([0.40, 0.30, 0.20, 0.10]) # SFT model's probabilities
reward = np.array([1.0, 1.5, 2.5, -1.0]) # reward model scores (it overrates fluff!)
def best_policy(beta):
# The policy that maximises E[r] - beta * KL(pi || pi_ref)
# has the closed form pi*(y) = pi_ref(y) * exp(r(y) / beta) / Z
w = pi_ref * np.exp(reward / beta)
return w / w.sum()
for beta in (10.0, 1.0, 0.1):
pi = best_policy(beta)
kl = np.sum(pi * np.log(pi / pi_ref))
print(f"beta={beta:<4} KL={kl:.2f} " +
" ".join(f"{name}:{p:.2f}" for name, p in zip(replies, pi)))
# Per-sample reward actually used inside PPO for one sampled reply
y = 1 # sampled "detailed correct"
pi = best_policy(1.0)
r_total = reward[y] - 1.0 * (np.log(pi[y]) - np.log(pi_ref[y]))
print(f"\nr_RM={reward[y]}, log-ratio={np.log(pi[y] / pi_ref[y]):.2f}, r_total={r_total:.2f}")Output:
beta=10.0 KL=0.00 short correct:0.39 detailed correct:0.31 flattering fluff:0.23 wrong but confident:0.08 beta=1.0 KL=0.28 short correct:0.22 detailed correct:0.27 flattering fluff:0.50 wrong but confident:0.01 beta=0.1 KL=1.61 short correct:0.00 detailed correct:0.00 flattering fluff:1.00 wrong but confident:0.00 r_RM=1.5, log-ratio=-0.09, r_total=1.59
Putting it all together
One full RLHF iteration for our support bot: sample 512 customer prompts, let the policy write replies, score them with the reward model, subtract β times the per-token log-ratio against the SFT model, estimate advantages with the value model, run a few PPO epochs, then repeat. Every so often, show humans fresh samples, collect new comparisons, and retrain the reward model on the policy’s current style of answers.
Reward hacking
Reward hacking (or reward over-optimisation) happens when the policy finds ways to raise the reward model’s score that do not reflect real quality. The reward model is a proxy for human judgement, and pushing hard on any proxy eventually breaks it (an instance of Goodhart’s law: when a measure becomes a target, it stops being a good measure).
- Length hacking: reward models often prefer longer answers, so the policy pads its replies.
- Sycophancy: agreeing with the user or flattering them scores well even when the user is wrong.
- Style over substance: confident tone, bullet lists or hedging phrases that labelers liked in training data get over-used.
- Degenerate text: with no KL penalty, odd repeated tokens can trigger high scores.
Pause and think: Average reply length doubled during RL training and reward scores rose. What should we check before celebrating?
Whether humans actually prefer the longer replies. Rising length with rising reward is a classic sign of length hacking. Compare with human ratings at matched length, consider a length penalty, raise β, or retrain the reward model with examples where shorter answers win.
Common mistakes and best practices
Common mistakes Skipping or rushing SFT, so RL starts from a weak policy; vague labeling guidelines that produce noisy, inconsistent comparisons; trusting rising reward-model scores without human evaluation; setting β too low (reward hacking) or too high (nothing changes); training the reward model once and never refreshing it as the policy changes; and evaluating only on the reward model you trained against.
- Invest in data quality: clear guidelines, trained labelers, agreement checks, and diverse prompts that match real use.
- Keep a held-out evaluation: human ratings or a separate judge model, never just the training reward model.
- Monitor KL, length and reward together: fast-rising KL or length is a warning sign.
- Tune β (or use an adaptive KL controller) to target a sensible KL budget.
- Refresh the reward model with comparisons on the current policy’s outputs.
- Consider simpler alternatives: DPO for offline preference data, or verifiable rewards (tests, exact answers) where they exist.
Where RLHF is used RLHF and its descendants are part of the post-training of most major chat assistants, typically alongside SFT and other preference methods. Exact recipes differ by organisation and are often only partly public. Outside chat, the same idea is used for summarisation, code assistants and image generation models tuned on human preferences.
Worked example, step by step
Let us follow one customer message through stages 2 and 3 with small numbers. The message: “My e-bike battery died after 20 minutes. What now?” All numbers are illustrative.
One prompt, from comparison to policy update
- A human compares two replies: Reply A gives the battery-swap steps. Reply B apologises three times and offers a refund we do not give. The labeler picks A.
- The untrained reward model disagrees: It scores A at 0.2 and B at 0.6. The loss is
−log σ(0.2 − 0.6) = −log σ(−0.4) ≈ −log 0.40 ≈ 0.91. That is above 0.69, the loss of a coin flip, because the model ranked the pair the wrong way. - The reward model learns: After training on many such pairs it scores A at 1.4 and B at 0.4. The loss for this pair falls to about 0.31. The reward model is now frozen.
- The policy writes a new reply: In stage 3 the policy writes reply C. The reward model gives it 1.2. Its summed log-ratio against the reference model is 3.0: the policy has made this reply somewhat more likely than the SFT model did.
- Subtract the KL penalty: With
β = 0.1:r_total = 1.2 − 0.1 × 3.0 = 0.9. - Compare with what was expected: The value model expected 0.5 for this prompt. The advantage is
0.9 − 0.5 = +0.4: better than expected. PPO raises the probability of reply C’s tokens, within its clip range.
Now a second reply to the same prompt, to see the penalty bite:
| Reply | RM score | Log-ratio | r_total | Advantage | Effect |
|---|---|---|---|---|---|
| C: steps to swap the battery | 1.2 | 3.0 | 0.9 | +0.4 | Made more likely |
| D: very long, warm, vague | 1.6 | 14.0 | 0.2 | −0.3 | Made less likely |
Reply D has the higher reward-model score, yet it is pushed down. It sits far from what the reference model would write, and the penalty of 0.1 × 14.0 = 1.4 outweighs its lead. This is the KL term doing its job: a high score earned far from familiar ground is treated with suspicion.
Practice: try it yourself
Earlier we computed where the policy ends up for a given β. Now we will watch it get there. We simulate stage 3 as a training loop over four possible replies. The reward model has one flaw: it overrates long, flattering replies. We track what the reward model sees and, since this is a simulation, the true quality it cannot see.
practice_rlhf_loop.py
import numpy as np
replies = ["short correct", "detailed correct", "long flattering", "wrong"]
pi_ref = np.array([0.40, 0.30, 0.20, 0.10]) # frozen SFT model
rm_score = np.array([1.0, 1.5, 2.5, -1.0]) # reward model (overrates flattery)
true_quality = np.array([1.0, 1.5, 0.2, -1.0]) # what people really think
beta = 0.5
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
logits = np.log(pi_ref) # the policy starts as the SFT copy
print("step RM reward true quality KL P(long flattering)")
for step in range(0, 201):
pi = softmax(logits)
kl = max(0.0, float(np.sum(pi * np.log(pi / pi_ref))))
if step in (0, 10, 50, 200):
print(f"{step:>4} {pi @ rm_score:9.2f} {pi @ true_quality:12.2f} {kl:5.2f} {pi[2]:18.2f}")
# Reward each reply gets: RM score minus the KL penalty term
r_total = rm_score - beta * np.log(pi / pi_ref)
advantage = r_total - pi @ r_total # better or worse than average?
logits += 0.2 * pi * advantage # policy-gradient step
print("final policy:", dict(zip(replies, softmax(logits).round(2).tolist())))Output:
step RM reward true quality KL P(long flattering)
0 1.25 0.79 0.00 0.20
10 1.50 0.77 0.05 0.32
50 2.07 0.50 0.57 0.68
200 2.22 0.44 0.79 0.77
final policy: {'short correct': 0.07, 'detailed correct': 0.15, 'long flattering': 0.77, 'wrong': 0.01}The reward-model column rises on every row. The true-quality column falls on every row. Nothing inside the training loop can see the second column. That gap is reward hacking.
Now change it:
- Raise
betaon line 7 from0.5to3.0. Predict: where does P(long flattering) settle, and does true quality end above or below its starting value of 0.79? - Repair the reward model: on line 5 change
2.5to0.2, so it matches true quality. Keepbeta = 0.5. Predict which reply the policy now favours and what happens to true quality. - Stop early: change
range(0, 201)on line 15 torange(0, 11). Predict the final policy. Is early stopping a substitute for a KL penalty, or a different tool?
Pause and think: KL settles near 0.79 and the policy stops changing, even though “long flattering” still has the highest reward-model score. What stops it?
The penalty has caught up. At the end the policy gives that reply 0.77 against the reference’s 0.20, so it pays β · log(0.77 / 0.20) ≈ 0.5 × 1.35 ≈ 0.67 every time. Its reward minus penalty is now equal to that of the other replies, so every advantage is zero and the update vanishes. A larger β reaches this balance sooner and closer to the reference.
Pause and think: At step 10 true quality is 0.77, almost unchanged. By step 50 it is 0.50. In a real project we cannot print true quality. How could we still notice the slide between those two checkpoints?
By evaluating saved checkpoints with something the policy was not trained against: fresh human ratings or a separate held-out judge. Inside training, warning signs are a fast-rising KL (0.05 → 0.57 here) and one style of reply taking over (0.32 → 0.68). The reward-model score itself is useless for this, because it is the very thing being exploited.
Quick summary
RLHF turns human comparisons into model improvements. Stage 1 (SFT) teaches the format and gives a good start. Stage 2 trains a reward model with −log σ(r_chosen − r_rejected). Stage 3 uses PPO to maximise r_RM − β·log(π/π_ref). The KL penalty is what keeps optimisation honest: it limits how far the policy can drift into regions where the reward model is wrong. Reward hacking is the main failure mode, so always judge results with fresh human or independent evaluation.
Key takeaways
- RLHF = SFT → reward model from human comparisons → RL (usually PPO).
- The reward model is trained with −log σ(r_chosen − r_rejected); only score differences matter.
- The KL penalty r_RM − β·log(π/π_ref) keeps the policy where the reward model is trustworthy.
- Reward hacking (length, sycophancy, style) is the main failure; watch reward, KL, length and human ratings together.
- Judge success with fresh human or independent evaluation, not the training reward model.
Key terms
- RLHF: Reinforcement learning from human feedback: optimising a model against a reward learned from human preferences.
- Reward model: A model that gives a scalar score predicting human preference for a response.
- Policy: In RLHF, the language model being trained, viewed as choosing tokens (actions).
- Reference model: A frozen copy of the SFT model used to compute the KL penalty.
- KL penalty: A cost proportional to how far the policy’s distribution moves from the reference model.
- Reward hacking: Raising the proxy reward without improving real quality.
- Value model (critic): A model PPO uses to estimate expected reward, so advantages can be computed.
← 8.7 InstructGPT: Teaching GPT-3 to Follow Instructions · 8.9 PPO: The Reinforcement Algorithm Behind Instruction Tuning →