Modern AI Engineering

Lesson 8.7 · 27 min

InstructGPT: Teaching GPT-3 to Follow Instructions

How did a 1.3-billion-parameter model end up preferred by people over the 175-billion-parameter GPT-3 it was built from?

In short: InstructGPT (OpenAI, 2022) turned GPT-3 from a text-continuation engine into a model that follows instructions. It used three steps: supervised fine-tuning on human-written answers, a reward model trained on human rankings of answers, and reinforcement learning with PPO against that reward, kept close to the original model by a KL penalty. People preferred its answers strongly, it made up facts less often, and the recipe became the basis of ChatGPT-style assistants.

What is the InstructGPT paper?

“Training language models to follow instructions with human feedback” by Long Ouyang and colleagues at OpenAI (2022) describes InstructGPT: GPT-3 models fine-tuned so that they do what the user asks, rather than just continuing the text. It was the first large-scale demonstration of RLHF (reinforcement learning from human feedback) on a general-purpose language model, and the same approach was used for ChatGPT later that year.

Think of it like coaching a brilliant but literal new hire The new hire has read every book in the library (pretraining) but answers a request like “write a short apology email” by writing three more requests, because that is what the documents they read looked like. Coaching has three stages: show them good examples, teach them to judge which drafts are better, then let them practise and reward the drafts that a judge likes.

The building blocks we must know first

  • Language model (LM): a model that predicts the next token. GPT-3 (2020) is a 175-billion-parameter LM trained on internet text.
  • Prompt: the input text we give the model. Completion / response: what it writes back.
  • Supervised fine-tuning (SFT): further training on input → ideal-output examples, with the normal next-token loss.
  • Reward model (RM): a model that reads a prompt and a response and outputs one number: how good humans would find it.
  • Reinforcement learning (RL): improving a policy (here, the LM) by trying actions (writing responses) and increasing the probability of those that earn high reward.
  • PPO (Proximal Policy Optimization): a popular, stable RL algorithm that limits how far each update can move the policy. It has its own lesson later.
  • KL divergence: a measure of how different two probability distributions are; used to keep the trained model close to a reference model.

Why GPT-3 was not enough

GPT-3 was trained to predict the next token on internet text. That objective is not the same as “help the user safely”. The paper calls this gap misalignment: the model optimises one thing (plausible continuation) while users want another (a helpful answer).

  • Not following instructions: asked “Explain the moon landing to a 6-year-old”, a base model might continue with a list of similar prompts, since that pattern appears online.
  • Making things up: it confidently produces false facts, because plausible text is rewarded, not true text.
  • Toxic or biased output: it can reproduce harmful text from its training data.
  • Needing careful prompt engineering: to get useful behaviour, users had to write few-shot examples or clever prompt formats.

Running example: a user asks our bike-rental assistant, “Write a two-sentence reply to a customer whose e-bike battery died mid-ride.” GPT-3 might continue with “Write a reply to a customer whose…”, ramble, or invent a refund policy. We want a short, kind, accurate reply.

Helpful, honest, and harmless

The paper frames alignment using three goals popularised by Askell and colleagues at Anthropic (2021):

  • Helpful: does what the user intends, including inferring intent from short or unclear instructions, and asks for clarification when needed.
  • Honest: does not make up information or mislead. Because we cannot read a model’s “beliefs”, the paper measured proxies: truthfulness on benchmarks and how often it hallucinated (invented facts).
  • Harmless: avoids causing physical, psychological or social harm, such as toxic or dangerous content.

These goals can conflict: the most “helpful” answer to a harmful request is not harmless. In InstructGPT’s training data, labelers were told to prioritise helpfulness to the user; in the final evaluations they were asked to weigh truthfulness and harmlessness more heavily. Defining these trade-offs is a judgement call that varies between organisations.

Pause and think: A user asks the assistant for a refund policy our shop does not have. Which of the three goals is most at risk if the model invents one?

Honest. Inventing a plausible-sounding policy is a hallucination. A helpful and honest reply says it does not know the policy and points the user to where to check.

The three-step method

The prompts came mostly from users of the OpenAI API (with personal information filtered out) plus prompts written by labelers to get started. A team of about 40 contract labelers wrote demonstrations and ranked outputs.

Step 1: supervised fine-tuning

Labelers wrote high-quality answers to prompts. GPT-3 was fine-tuned on these prompt → answer pairs with the usual next-token cross-entropy loss. This teaches the format of following instructions: answer the question, do not continue the list.

An interesting detail: the authors trained SFT for 16 epochs, which overfit by validation loss, but kept going because further training still improved reward-model scores and human preference ratings. This is a reminder that the loss we optimise and the quality we want are not the same thing.

Why not stop here? Writing ideal answers is slow and expensive, and SFT only imitates. It cannot learn that answer A is better than answer B unless someone writes the better one. Ranking is cheaper and gives a richer signal.

Step 2: the reward model

For each prompt, the model generated K answers (K between 4 and 9), and a labeler ranked them from best to worst. A ranking of K answers contains K·(K−1)/2 pairwise comparisons: 6 for K = 4, 36 for K = 9. Ranking 9 answers once is much faster than judging 36 separate pairs.

The reward model was a 6-billion-parameter model initialised from the SFT model, with the final next-token layer replaced by a layer that outputs one number. (They reported that a 175B reward model was less stable to train.) The loss uses the Bradley–Terry idea from the previous lesson: for each pair where answer y_w beat y_l, push the score of the winner above the loser.

One practical trick: the authors put all C(K,2) comparisons from one prompt into the same batch element. Treating them as separate shuffled examples made the reward model overfit, because each answer appears in many correlated pairs.

ranking_to_pairs.py

import numpy as np
from itertools import combinations
# A labeler ranked K=4 answers to one prompt, best first: C > A > D > B
ranking = ["C", "A", "D", "B"]
# Scalar scores the reward model currently gives each answer
rm_score = {"A": 1.2, "B": -0.3, "C": 0.9, "D": 0.1}
def log_sigmoid(x):
return -np.log1p(np.exp(-x))
pairs = list(combinations(ranking, 2))      # every (winner, loser) pair
print(f"K=4 ranking gives {len(pairs)} comparisons")
losses = []
for win, lose in pairs:                     # earlier in ranking = winner
diff = rm_score[win] - rm_score[lose]
loss = -log_sigmoid(diff)               # -log sigma(r_w - r_l)
losses.append(loss)
flag = "  <- model has this pair backwards" if diff < 0 else ""
print(f"{win} > {lose}: r_w - r_l = {diff:+.1f}  loss = {loss:.3f}{flag}")
# InstructGPT averages over all C(K,2) pairs of one prompt together
print(f"mean loss for this prompt = {np.mean(losses):.3f}")

Output:

K=4 ranking gives 6 comparisons
C > A: r_w - r_l = -0.3  loss = 0.854  <- model has this pair backwards
C > D: r_w - r_l = +0.8  loss = 0.371
C > B: r_w - r_l = +1.2  loss = 0.263
A > D: r_w - r_l = +1.1  loss = 0.287
A > B: r_w - r_l = +1.5  loss = 0.201
D > B: r_w - r_l = +0.4  loss = 0.513
mean loss for this prompt = 0.415

Pause and think: A labeler ranks 6 answers to one prompt. How many pairwise comparisons does the reward model learn from?

C(6,2) = 6·5/2 = 15 comparisons, all from one ranking.

Step 3: reinforcement learning with PPO

Now the SFT model becomes the policy. For each prompt from a fresh set, it writes an answer; the reward model scores it; PPO updates the policy to make high-scoring answers more likely. Each whole answer is treated as one action with one reward at the end (a “bandit” setting).

The KL penalty matters a lot. Without it the policy can find odd answers that the reward model scores highly but humans dislike (reward hacking), or drift into repetitive text. Subtracting β · log(π_RL / π_SFT) charges the policy for every token where it strays from the SFT model.

One PPO round in InstructGPT

  1. Sample prompts: Take a batch of prompts from the PPO prompt set.
  2. Generate answers: The current policy writes one answer per prompt.
  3. Score: The reward model gives each answer a number; the KL penalty against the SFT model is subtracted.
  4. Update with PPO: Increase the probability of tokens in answers that beat expectations, with PPO’s clipping to keep each update small.
  5. Mix pretraining (PPO-ptx): Also take gradient steps on ordinary pretraining text, to protect general abilities.

The alignment tax and the results

The alignment tax is the drop in performance on some standard tasks that comes from alignment training. After plain PPO, InstructGPT got worse on several public NLP benchmarks (such as SQuAD, DROP, HellaSwag and WMT French→English translation). PPO-ptx, which mixes in pretraining gradients, greatly reduced these regressions without hurting human preference scores much. That is why PPO-ptx became the main InstructGPT model.

  • Preference: outputs from the 1.3B InstructGPT were preferred over the 175B GPT-3, despite having over 100× fewer parameters.
  • Truthfulness: on the TruthfulQA benchmark, InstructGPT produced truthful and informative answers about twice as often as GPT-3, and it made up facts less often on closed-domain tasks such as summarisation.
  • Toxicity: when asked to be respectful, it produced about 25% fewer toxic outputs than GPT-3; with no such instruction, the gain largely disappeared.
  • Bias: no significant improvement on the bias benchmarks they tested.
  • Generalisation: it followed instructions somewhat in areas rare in the fine-tuning data, such as code and non-English prompts, and was preferred by held-out labelers who had produced no training data.
  • Still imperfect: it could still follow harmful instructions, make simple mistakes, and hedge too much.

Common misconception InstructGPT is not aligned “with humanity”. It is aligned with the preferences of a specific group: about 40 labelers, the instructions they were given by the researchers, and the API customers whose prompts were used. The paper itself stresses this.

Going one level deeper

The step-3 objective is easier to trust once we have put numbers into it. Take one prompt from our bike-rental assistant and three answers the policy might write. For each we need two numbers: the reward model’s score r, and the log-ratio log(π_RL / π_SFT), which says how much more likely the current policy makes this answer than the SFT model did. We use β = 0.02 here; all numbers are illustrative.

Reward minus KL penalty for three candidate answers (illustrative, β = 0.02)
AnswerReward rLog-ratioPenalty β · log-ratioObjective
A: plain, in the SFT style1.220.041.16
B: clearer and more specific2.0150.301.70
C: odd text the reward model happens to love3.01202.400.60

Reading the table

  1. Where a log-ratio of 15 comes from: The log-ratio is summed over tokens. If answer B has 30 tokens and the policy makes each one about 0.5 more likely in log terms, the total is 30 × 0.5 = 15.
  2. Without the penalty: With β = 0 the objective is just the reward, so C (3.0) wins. The policy would be pulled toward text that the SFT model would almost never write.
  3. With the penalty: C pays 0.02 × 120 = 2.40 and drops to 0.60. B pays only 0.30 and leads with 1.70. The policy is pulled toward B: better, and still close to home.
  4. Why “close to home” matters: The reward model was trained on answers that look like SFT output. Its score for C is a guess far outside that range, and such guesses are the least reliable ones.
  5. The third term: PPO-ptx adds γ times the log-likelihood of ordinary pretraining text. It does not look at these answers at all. It asks a separate question: can the model still predict normal text well?

Two things follow. First, the penalty grows with length, since it is a sum over tokens; a long answer that drifts a little per token can pay as much as a short one that drifts a lot. Second, β is a dial, not a constant of nature: larger β keeps the policy nearer the SFT model and improves less; smaller β improves more on the reward model’s terms and risks answers like C.

Practice: try it yourself

We will train the smallest possible reward model: one score per answer. Three labelers rank the same four answers and do not fully agree. We turn their rankings into pairs, run gradient descent on the pairwise loss from step 2, and see what scores come out when humans disagree.

practice_reward_scores.py

import numpy as np
from itertools import combinations
answers = ["A", "B", "C", "D"]
# Three labelers rank the same 4 answers, best first. They do not fully agree.
rankings = [["A", "B", "C", "D"],
["A", "C", "B", "D"],
["B", "A", "C", "D"]]
# Every ranking of K = 4 answers gives C(4,2) = 6 (winner, loser) pairs.
pairs = [p for rank in rankings for p in combinations(rank, 2)]
print("total pairs:", len(pairs))
scores = {a: 0.0 for a in answers}        # the reward model's score per answer
lr = 0.5
for step in range(200):
grads = {a: 0.0 for a in answers}
for win, lose in pairs:
p = 1 / (1 + np.exp(-(scores[win] - scores[lose])))   # sigma(r_w - r_l)
grads[win] -= (1 - p)             # gradient of -log(p): winner goes up
grads[lose] += (1 - p)            # and the loser goes down
for a in answers:
scores[a] -= lr * grads[a] / len(pairs)
for a in answers:
print(f"score {a}: {scores[a]:+.2f}")
sig = lambda x, y: 1 / (1 + np.exp(-(scores[x] - scores[y])))
print(f"P(A beats B) = {sig('A', 'B'):.2f}   (2 of 3 labelers agreed)")
print(f"P(B beats C) = {sig('B', 'C'):.2f}   (2 of 3 labelers agreed)")
print(f"P(C beats D) = {sig('C', 'D'):.2f}   (3 of 3 labelers agreed)")

Output:

total pairs: 18
score A: +2.20
score B: +1.10
score C: +0.02
score D: -3.31
P(A beats B) = 0.75   (2 of 3 labelers agreed)
P(B beats C) = 0.75   (2 of 3 labelers agreed)
P(C beats D) = 0.97   (3 of 3 labelers agreed)

The scores keep the majority order A > B > C > D, and the gaps reflect how consistently the labelers agreed: A over B is 0.75, while C over D, where everyone agreed, is 0.97.

Now change it:

  • Change the third ranking on line 8 to ["A", "B", "C", "D"], so A always beats B. Predict: does P(A beats B) rise a little, or head toward 1.00?
  • Change range(200) on line 16 to range(2000). Predict: which answer’s score moves the most with the extra training, and why does it never settle?
  • Add a fourth labeler with the reversed ranking ["D", "C", "B", "A"]. Predict: do the scores spread out or move closer together?

Pause and think: Answer D ends with a score of −3.31. Does that mean D is a harmful or terrible answer?

Not necessarily. The loss only uses score differences, so −3.31 means “D lost every comparison against A, B and C”. All four answers could be good, with D simply the least good of this set. A reward model’s number has meaning only relative to other answers, which is why the paper shifts the reward model’s output to a fixed zero point before the RL step.

Pause and think: The labelers put A above B in 2 of 3 rankings (67%), yet the model says P(A beats B) = 0.75. Why is it not exactly 0.67?

Each answer gets a single score, and that score has to explain all its comparisons at once. A also beat C in 3 of 3 rankings while B beat C in only 2 of 3, which is extra evidence that A sits above B. The fit is a compromise over all 18 pairs, not a copy of each pair’s win rate. This sharing is also what lets a reward model score answers it has never seen compared.

What alignment looks like today, and a quick summary

The three-step recipe (SFT, reward model, RL) shaped the first wave of chat assistants. Since then, teams have added and swapped parts: DPO and similar methods learn directly from preference pairs without a separate reward model or PPO loop; RLAIF and constitution-style methods use AI feedback to scale labelling; and reinforcement learning with verifiable rewards (for example checking math answers or running code tests), often with GRPO, drives reasoning models. Exact pipelines vary by lab and are often only partly published.

Quick summary: InstructGPT showed that GPT-3’s problem was misalignment, not lack of knowledge. Fine-tune on demonstrations (SFT), learn a reward model from rankings (K answers give C(K,2) pairs, loss −log σ(r_w − r_l)), and optimise with PPO plus a KL penalty and pretraining mix (PPO-ptx). The result: a 1.3B model people preferred over 175B GPT-3, more truthful and less toxic when asked, at a small and mostly recoverable alignment tax.

Key takeaways

  • GPT-3’s problem was misalignment: next-token prediction is not the same as helping the user.
  • InstructGPT = SFT on demonstrations → reward model on rankings → PPO with a KL penalty.
  • Rankings of K answers give C(K,2) comparisons, trained with −log σ(r_w − r_l).
  • PPO-ptx mixes in pretraining gradients to reduce the alignment tax.
  • A 1.3B aligned model was preferred over 175B GPT-3: human preference data can beat raw scale.

Key terms

  • InstructGPT: GPT-3 models fine-tuned with SFT and RLHF to follow instructions (OpenAI, 2022).
  • Alignment: Making a model’s behaviour match what its users and developers intend.
  • SFT: Supervised fine-tuning on human-written demonstration answers.
  • Reward model: A model that outputs a scalar score predicting how much humans would like a response.
  • KL penalty: A cost for the policy’s distribution moving away from the reference (SFT) model.
  • Alignment tax: Performance lost on some tasks as a side effect of alignment training.
  • PPO-ptx: PPO training mixed with pretraining-data gradients to limit the alignment tax.

← 8.6 Deep RL from Human Preferences: The Foundational Paper · 8.8 RLHF: Aligning LLMs with Human Preferences →