Lesson 8.11 · 25 min
GRPO: Group-Based Preference Optimization Explained
Instead of training a whole second network to guess how good an answer “should” be, what if we just asked the model the same question eight times and compared the answers with each other?
In short: GRPO (Group Relative Policy Optimization) is a reinforcement learning method for language models that removes PPO’s value model. For each prompt it samples a group of answers, scores them, and uses each answer’s reward relative to the group’s mean (divided by the group’s standard deviation) as its advantage. It keeps PPO’s clipped update and a KL penalty to a reference model. Introduced with DeepSeekMath and used for DeepSeek-R1, it is a popular choice for training reasoning models with verifiable rewards.
What is GRPO?
GRPO, short for Group Relative Policy Optimization, was introduced by Zhihong Shao and colleagues at DeepSeek in the 2024 DeepSeekMath paper. It became widely known in 2025 when DeepSeek-R1, a model with strong step-by-step reasoning, was trained with large-scale reinforcement learning using GRPO.
GRPO is a variant of PPO (Proximal Policy Optimization). It keeps PPO’s core, the clipped probability-ratio update, but changes how the advantage (how much better an answer was than expected) is computed: instead of a learned value model, it compares answers within a group sampled for the same prompt.
Think of it like grading on a curve A teacher without an answer key for “how hard was this exam?” can still grade fairly by comparing students who took the same exam: above the class average is good, below is bad. GRPO grades each answer against its classmates, the other answers to the same prompt, instead of against a separately trained predictor of expected score.
Why do we need GRPO? The problem with PPO
In PPO for language models we need an advantage for every token: A = actual reward − expected reward. The expected part comes from a value model (or critic), a neural network usually as large as the policy, trained alongside it.
- Memory and compute: the value model is a second large network to store, run and train. Together with the policy, reference and reward model, that is four big models.
- Hard to train well: the reward usually arrives only at the end of the answer, but the value model must predict it at every token. That is a difficult regression problem, and a poor value model gives noisy advantages.
- Extra hyperparameters: value loss weight, GAE settings and more, each a source of instability.
Running example for this lesson: we train a model for our bike-rental shop that answers pricing questions requiring arithmetic, such as “3 days of e-bike at 18 per day with a 10% weekly-pass discount: what is the total?” A program can check the final number, so the reward is simply 1 if correct and 0 if not.
Pause and think: If the reward is 1 or 0 only at the end of an answer, why is a value model’s job hard?
It must estimate, from a half-written answer, the probability that the final number will be right. Early tokens give little evidence, and the target is noisy (0 or 1). Errors in this estimate become errors in every token’s advantage.
How does GRPO work?
One GRPO step for one prompt
- Sample a group: Generate G answers (for example 8) to the same prompt from the current policy, with some randomness (temperature above 0).
- Score each answer: Use a reward function: a rule-based checker (correct final answer, valid format, passing unit tests) or a reward model.
- Compute group-relative advantages: Aᵢ = (rᵢ − mean(r)) / std(r). Answers above the group average get positive advantages; below average get negative.
- Share it across tokens: Every token of answer i gets the same advantage Aᵢ (with outcome rewards).
- Clipped update plus KL: Apply PPO’s clipped ratio objective with these advantages, and subtract a KL penalty that keeps the policy close to a frozen reference model.
Step-by-step example
We ask the pricing question 8 times. A checker finds 3 correct answers (reward 1) and 5 wrong ones (reward 0). The group mean is 3/8 = 0.375 and the standard deviation is about 0.484. So each correct answer gets (1 − 0.375) / 0.484 ≈ +1.29 and each wrong one gets (0 − 0.375) / 0.484 ≈ −0.77. The code computes this, then applies the PPO-style clipped term and a KL estimate.
grpo_advantages.py
import numpy as np
# One math prompt, a group of G = 8 sampled answers, checked by a rule:
# reward 1 if the final answer is correct, else 0
rewards = np.array([1, 0, 0, 1, 0, 0, 0, 1], dtype=float)
mean, std = rewards.mean(), rewards.std()
adv = (rewards - mean) / (std + 1e-8) # group-relative advantage
print(f"group mean {mean:.3f}, std {std:.3f}")
for i, (r, a) in enumerate(zip(rewards, adv)):
print(f"answer {i}: reward {r:.0f} -> advantage {a:+.2f}")
# Every token of answer i shares advantage adv[i]; PPO-style clipped term
eps = 0.2
ratio = np.array([1.05, 0.97, 1.30, 1.10, 0.92, 1.00, 0.70, 1.25]) # new/old prob
surrogate = np.minimum(ratio * adv, np.clip(ratio, 1 - eps, 1 + eps) * adv)
print("clipped terms:", np.round(surrogate, 2))
# KL estimate to the reference model, used by GRPO (always >= 0)
logp, logp_ref = -1.20, -1.50 # one token's log-probs
x = np.exp(logp_ref - logp)
print(f"k3 KL estimate = {x - np.log(x) - 1:.4f}")
all_same = np.ones(8) # every answer correct
print("all-correct group advantages:", (all_same - all_same.mean()) / (all_same.std() + 1e-8))Output:
group mean 0.375, std 0.484 answer 0: reward 1 -> advantage +1.29 answer 1: reward 0 -> advantage -0.77 answer 2: reward 0 -> advantage -0.77 answer 3: reward 1 -> advantage +1.29 answer 4: reward 0 -> advantage -0.77 answer 5: reward 0 -> advantage -0.77 answer 6: reward 0 -> advantage -0.77 answer 7: reward 1 -> advantage +1.29 clipped terms: [ 1.36 -0.75 -1.01 1.42 -0.71 -0.77 -0.62 1.55] k3 KL estimate = 0.0408 all-correct group advantages: [0. 0. 0. 0. 0. 0. 0. 0.]
Pause and think: Suppose only 1 of 8 answers is correct. Will its advantage be larger or smaller than +1.29? Why?
Larger. Mean = 0.125 and std ≈ 0.331, so the correct answer gets (1 − 0.125)/0.331 ≈ +2.65 and wrong ones ≈ −0.38. A rare success on a hard prompt is a strong signal, so it gets a big push; the many failures are each pushed down only a little.
The GRPO objective in simple words
- Average over the group and over tokens: every answer contributes; inside each answer, every token gets the answer’s advantage.
- PPO clipping: each token’s probability can only change by about ±ε per round, keeping updates stable.
- KL in the loss, not in the reward: PPO-style RLHF usually subtracts the KL penalty from the reward; GRPO adds it as a separate term in the loss, so it does not muddy the advantage calculation.
PPO vs GRPO
Advantages of GRPO
- Memory savings: removing the value model frees a large share of training memory, which can go to bigger batches or longer answers.
- Simplicity: fewer models and fewer hyperparameters than PPO.
- Natural fit for comparisons: reward models are trained on comparisons between answers to the same prompt, and GRPO also compares answers to the same prompt.
- Strong with verifiable rewards: with simple rule-based rewards (correct answer, correct format), DeepSeek reported that reasoning behaviours such as long step-by-step thinking and self-checking emerged during RL training of DeepSeek-R1-Zero.
Where GRPO is used DeepSeekMath and DeepSeek-R1 used GRPO. Since then it has been implemented in open-source training libraries such as Hugging Face TRL, and many open reasoning-model projects use GRPO or close variants for math, coding and other tasks where a program can check the answer.
Practical things to keep in mind
- Group size: larger groups give steadier baselines but cost more generation. Values from about 4 up to 64 appear in practice.
- Zero-variance groups: if all answers are right (too easy) or all wrong (too hard), every advantage is 0 and the prompt is wasted. Choose prompts at the right difficulty, or filter such groups.
- Reward design: rule-based rewards are hard to hack but can still be gamed, for example by guessing formats the checker accepts. Test the checker on tricky cases.
- Sampling temperature: answers in a group must differ, or there is nothing to compare. Too low a temperature removes diversity.
- Watch length and KL: reasoning RL often makes answers longer; that can be good (more thinking) or a bias. Track length, KL to the reference, and accuracy on held-out problems.
Common mistakes Using a prompt set that is far too easy or too hard (most groups have zero variance); trusting a buggy answer checker; setting the temperature so low that all G answers are identical; and assuming GRPO needs no reference model (most implementations keep one for the KL term, although some recipes drop it).
Going one level deeper
Our worked example used one prompt. Real training batches mix easy and hard prompts, and this is where “relative to the group” starts to matter. Take two pricing questions with 4 answers each. On the easy one, three answers are correct. On the hard one, only one is. Rewards are 1 for correct and 0 for wrong.
| Prompt | Rewards | Mean | Std | Advantage of a correct answer | Advantage of a wrong answer |
|---|---|---|---|---|---|
| Easy | [1, 1, 1, 0] | 0.75 | 0.433 | +0.58 | −1.73 |
| Hard | [0, 0, 0, 1] | 0.25 | 0.433 | +1.73 | −0.58 |
Reading the two rows
- Easy prompt, correct answer:
(1 − 0.75) / 0.433 ≈ +0.58. Being right here is normal, so it earns a small push up. - Easy prompt, wrong answer:
(0 − 0.75) / 0.433 ≈ −1.73. Failing where most samples succeed is a strong signal: push this answer down hard. - Hard prompt, correct answer:
(1 − 0.25) / 0.433 ≈ +1.73. A rare success is the most valuable sample in the batch. - Hard prompt, wrong answers:
−0.58each. Failing a hard question is expected, so the push down is mild. - Compare with raw rewards: With raw rewards, every correct answer would get +1 and every wrong one 0. Wrong answers would never be pushed down, and the rare success on the hard prompt would count no more than a routine one.
Dividing by the standard deviation has a second effect that cuts both ways. It makes the advantage independent of reward scale: rewards [10, 0, 0, 0] give exactly the same advantages as [1, 0, 0, 0]. That is convenient when prompts use different scoring ranges.
The same property is also a failure case. Suppose a graded reward gives [0.51, 0.50, 0.50, 0.50], a difference that is probably noise. After normalisation the first answer still gets +1.73, as if it were a clear winner. So when rewards are graded rather than 0/1, we should make sure small differences are meaningful, or round the scores before training.
Practice: try it yourself
We will run 40 GRPO steps on the smallest policy we can write: a single number theta that sets the chance of answering one pricing question correctly. Each step samples a group of 8 answers, checks them, computes group-relative advantages and updates theta. We also count how many groups carried no signal at all.
practice_grpo_loop.py
import numpy as np
rng = np.random.default_rng(0)
G = 8 # answers sampled per prompt
theta = -1.5 # one policy parameter: P(correct answer) = sigmoid(theta)
lr = 0.5
skipped = 0
def p_correct(theta):
return 1 / (1 + np.exp(-theta))
print("step P(correct) rewards in the group")
for step in range(1, 41):
p = p_correct(theta)
rewards = (rng.random(G) < p).astype(float) # 1 = checker says correct
if step in (1, 10, 20, 30, 40):
print(f"{step:>4} {p:10.2f} {rewards.astype(int)}")
if rewards.std() == 0: # all right or all wrong
skipped += 1 # no signal: nothing to learn
continue
adv = (rewards - rewards.mean()) / rewards.std() # group-relative advantage
# Gradient of log-probability for this toy policy:
# a correct answer gives (1 - p), a wrong answer gives (-p).
grad_logp = np.where(rewards == 1, 1 - p, -p)
theta += lr * np.mean(adv * grad_logp) # raise good, lower bad
print(f"final P(correct) = {p_correct(theta):.2f}")
print(f"groups with zero variance (skipped): {skipped} of 40")Output:
step P(correct) rewards in the group 1 0.18 [0 0 1 1 0 0 0 0] 10 0.57 [1 0 1 1 0 0 0 1] 20 0.90 [1 1 1 1 0 1 1 1] 30 0.96 [1 1 1 1 1 1 1 0] 40 0.98 [1 1 1 1 1 1 1 1] final P(correct) = 0.98 groups with zero variance (skipped): 13 of 40
The policy climbs from 18% to 98% correct without any value model. Notice that 13 of the 40 groups were skipped: once the policy is nearly always right, most groups are all-correct and teach nothing.
Now change it:
- Make the prompt far too hard: set
theta = -5.0on line 5 (under 1% correct). Predict the number of skipped groups and the final P(correct). - Shrink the group: set
G = 2on line 4. Predict: do more or fewer groups get skipped, and why does a tiny group so often have zero variance? - Remove the division by the standard deviation on line 21, leaving
adv = rewards - rewards.mean(). Predict: does learning get faster or slower? (Hint: for 0/1 rewards the standard deviation is at most 0.5.)
Pause and think: Groups are skipped at the start (all wrong) and at the end (all right). Both have zero variance. Are they equally worrying?
No. All-right groups at the end mean the prompt is mastered; skipping them costs nothing. All-wrong groups at the start mean the model never produces a success to learn from, and if that persists, training stalls completely (try theta = -5.0). The fix for the second kind is easier prompts first, larger groups, or partial-credit rewards so that some answers score above others.
Pause and think: This toy policy serves one prompt. A real model uses the same weights for thousands of prompts of mixed difficulty. Why does normalising within each prompt’s group matter more there?
Because one update mixes signals from all prompts. With raw rewards, easy prompts would dominate: they produce many 1s, so their answers get pushed up constantly, while hard prompts contribute almost nothing. Per-group normalisation puts every prompt on the same footing: each asks only “which of my answers were better than my average?”, so a rare success on a hard prompt counts strongly.
When to use GRPO, and conclusion
| Situation | Good choice | Why |
|---|---|---|
| Answers can be checked by a program (math, code tests) | GRPO | Cheap, reliable rewards; group comparison works well |
| Fixed dataset of human preference pairs | DPO | Offline and simple; no sampling needed |
| Learned reward model, dense rewards, large infrastructure | PPO | Value model gives per-token credit |
| Good demonstrations, no reward signal | SFT | Imitation is enough |
When not to use GRPO: if we cannot afford several generations per prompt, if rewards are nearly always the same for every answer, or if we only have a fixed offline preference dataset (DPO is simpler there).
Conclusion: GRPO is PPO without the critic. For each prompt it samples a group, scores each answer, and uses the group-normalised reward as the advantage, then applies PPO’s clipped update and a separate KL penalty. That makes RL for language models lighter and simpler, and it pairs especially well with verifiable rewards, which is why it became central to training reasoning models.
Key takeaways
- GRPO is PPO without a value model: the group of answers to the same prompt provides the baseline.
- Advantage = (reward − group mean) / group std, shared by every token of that answer.
- It keeps PPO’s clipped ratio update and adds a KL penalty to a reference model in the loss.
- Groups where every answer scores the same give no signal; pick prompts of the right difficulty.
- It fits verifiable rewards (math, code) and was central to DeepSeek-R1’s reasoning training.
Key terms
- GRPO: Group Relative Policy Optimization: PPO-style RL that uses group-normalised rewards instead of a value model.
- Group: Several answers sampled from the policy for the same prompt.
- Group-relative advantage: (reward − group mean) / group standard deviation.
- Value model (critic): A network PPO uses to predict expected reward; GRPO removes it.
- Verifiable reward: A reward computed by a program, such as checking a final answer or running tests.
- Reference model: A frozen copy of the starting policy used for the KL penalty.
← 8.10 DPO: Alignment Without the Separate Reward Model · 9.1 Chain-of-Thought Prompting: Making Models Reason Step by Step →