Modern AI Engineering

Lesson 5.2 · 23 min

Nucleus Sampling: Top-k and Top-p Demystified

Always picking the most likely word makes a chatbot dull and repetitive, but picking from all 100,000 words lets in nonsense. How do we keep the good choices and cut the junk?

In short: Top-k and top-p are filters applied to the next-token probabilities before sampling. Top-k keeps a fixed number of the most likely tokens; top-p (nucleus sampling) keeps the smallest set of top tokens whose probabilities add up to at least p. Both throw away the long tail of unlikely tokens and renormalize the rest. Top-p adapts to how confident the model is, which is why it is the more common default; both are usually combined with temperature.

How an LLM picks the next token

At each step an LLM outputs a logit (a raw score) for every token in its vocabulary. Temperature divides those logits, softmax turns them into probabilities that sum to 1, and then a decoding strategy chooses one token. The decoding strategy is what this lesson is about.

Running example: our travel-booking assistant is completing "For our trip we booked a hotel in ...". Suppose the model's probabilities for the top candidates are:

Think of it like a talent show shortlist A judge does not pick the winner from every person in the city, nor does she always pick the same favourite. She first makes a shortlist of strong acts, then chooses among them, more often the stronger ones. Top-k makes a shortlist of a fixed size ("the top 5"). Top-p makes a shortlist big enough to cover most of the talent ("whoever together accounts for 90% of the votes").

The problem with always picking the best token

Greedy decoding always takes the single most likely token. It is simple and predictable, but it has two problems for open-ended text:

  • Dull and repetitive. Human writing often uses words that are not the single most likely choice. Always taking the top token produces flat text, and in long outputs the model can fall into loops, repeating the same phrase because it keeps being the most likely continuation.
  • No variety. Every run gives the same answer. That is bad for brainstorming, for generating several drafts, or for any feature that should feel natural rather than robotic.

Greedy is still a good choice when there is one right answer (classification, extraction, many code tasks). For everything else we want sampling: picking at random according to the probabilities.

The problem with picking from every token

Pure sampling from the full distribution has the opposite problem. Each tail token is unlikely, but there are tens of thousands of them. Together they can hold a noticeable share of the probability. If 3% of the mass sits in nonsense tokens like "banana", then roughly once every 33 tokens the model picks something odd. In a 500-token answer that happens many times, and each odd token sends the following text further off track because the model conditions on its own mistakes.

The fix is to truncate the tail: remove unlikely tokens, keep the plausible ones, and sample among those. Top-k and top-p are two ways to decide where to cut.

Pause and think: If each of 50,000 tail tokens has probability 0.000001, how much total probability does the tail hold?

50,000 × 0.000001 = 0.05, or 5%. Individually negligible tokens add up, so roughly 1 in 20 sampled tokens would come from the tail.

What is top-k sampling? A step-by-step example

Top-k sampling keeps only the k tokens with the highest probabilities, sets all others to zero, renormalizes the kept ones so they sum to 1 again, and samples from them. k is a whole number such as 40 or 50 in practice; we use k = 3 for clarity.

Top-k with k = 3

  1. Sort: Order tokens by probability: Paris 0.50, Lyon 0.15, France 0.12, the 0.08, Nice 0.06, ...
  2. Keep the top k: Keep Paris, Lyon and France. Their total is 0.50 + 0.15 + 0.12 = 0.77.
  3. Zero the rest: "the", "Nice", "a", "Rome" and "banana" can no longer be chosen.
  4. Renormalize: Divide each kept probability by 0.77: Paris 0.65, Lyon 0.19, France 0.16.
  5. Sample: Draw one of the three at random with those probabilities.

The problem with top-k sampling

The weakness of top-k is that k is fixed, but the model's confidence is not. The shape of the distribution changes at every step:

  • When the model is confident ("The capital of France is" → Paris 92%), k = 3 still keeps two weak options. Sampling occasionally picks one of them, which can introduce a wrong fact.
  • When the model is uncertain ("My favourite food is" → many reasonable answers of similar probability), k = 3 cuts off perfectly good options and makes the text less diverse than it should be.

No single k is right for both situations. That observation motivated top-p.

What is top-p sampling? A step-by-step example

Top-p sampling, also called nucleus sampling, was proposed in the 2019 paper "The Curious Case of Neural Text Degeneration" by Holtzman and colleagues. Instead of a fixed count, it keeps the smallest set of most-likely tokens whose probabilities add up to at least p (for example p = 0.9). That set is the "nucleus". Its size changes automatically with the model's confidence.

Top-p with p = 0.8

  1. Sort: Paris 0.50, Lyon 0.15, France 0.12, the 0.08, Nice 0.06, ...
  2. Running total: 0.50 → 0.65 → 0.77 → 0.85. The total first reaches 0.8 after the 4th token.
  3. Keep the nucleus: Keep Paris, Lyon, France and "the" (total 0.85). Drop the rest.
  4. Renormalize: Divide by 0.85: Paris 0.59, Lyon 0.18, France 0.14, the 0.09.
  5. Sample: Draw one of these four.

top_k_top_p.py

import numpy as np
tokens = ["Paris", "Lyon", "France", "the", "Nice", "a", "Rome", "banana"]
probs = np.array([0.50, 0.15, 0.12, 0.08, 0.06, 0.05, 0.03, 0.01])
def top_k(p, k):
keep = np.argsort(p)[::-1][:k]          # indices of the k largest
out = np.zeros_like(p)
out[keep] = p[keep]
return out / out.sum()                  # renormalise to sum to 1
def top_p(p, threshold):
order = np.argsort(p)[::-1]             # sort high -> low
cum = np.cumsum(p[order])
n = np.searchsorted(cum, threshold) + 1 # smallest set reaching threshold
keep = order[:n]
out = np.zeros_like(p)
out[keep] = p[keep]
return out / out.sum()
def show(name, p):
kept = [f"{t}:{v:.2f}" for t, v in zip(tokens, p) if v > 0]
print(f"{name:<12} kept {len(kept)} -> " + " ".join(kept))
show("top-k k=3", top_k(probs, 3))
show("top-p p=0.8", top_p(probs, 0.8))
show("top-p p=0.9", top_p(probs, 0.9))
# A confident distribution: top-p shrinks, top-k does not
sure = np.array([0.92, 0.03, 0.02, 0.01, 0.01, 0.005, 0.003, 0.002])
show("sure, k=3", top_k(sure, 3))
show("sure, p=0.9", top_p(sure, 0.9))
# A flat (uncertain) distribution: top-p grows, top-k does not
flat = np.array([0.16, 0.15, 0.14, 0.13, 0.12, 0.11, 0.10, 0.09])
show("flat, k=3", top_k(flat, 3))
show("flat, p=0.9", top_p(flat, 0.9))

Output:

top-k k=3    kept 3 -> Paris:0.65 Lyon:0.19 France:0.16
top-p p=0.8  kept 4 -> Paris:0.59 Lyon:0.18 France:0.14 the:0.09
top-p p=0.9  kept 5 -> Paris:0.55 Lyon:0.16 France:0.13 the:0.09 Nice:0.07
sure, k=3    kept 3 -> Paris:0.95 Lyon:0.03 France:0.02
sure, p=0.9  kept 1 -> Paris:1.00
flat, k=3    kept 3 -> Paris:0.36 Lyon:0.33 France:0.31
flat, p=0.9  kept 7 -> Paris:0.18 Lyon:0.16 France:0.15 the:0.14 Nice:0.13 a:0.12 Rome:0.11

Pause and think: With probabilities [0.6, 0.25, 0.1, 0.05] and p = 0.9, how many tokens does top-p keep?

Running total: 0.6, 0.85, 0.95. It first reaches 0.9 at the third token, so 3 tokens are kept (then renormalized by dividing by 0.95).

Top-k vs top-p sampling

How top-k and top-p work with temperature

Temperature and these filters do different jobs. Temperature reshapes the distribution (sharper or flatter). Top-k and top-p truncate it (remove the tail). They are commonly applied in this order:

The order matters: temperature is applied before top-p, so a high temperature flattens the distribution and therefore makes the nucleus larger (more tokens are needed to reach p). Libraries mostly follow this order, but details can differ, so check the one you use.

Other truncation methods exist Newer methods such as min-p sampling keep tokens whose probability is at least some fraction of the top token's probability. They share the same goal: cut the tail in a way that adapts to the model's confidence. Support varies by library and API.

When to use which one, and common mistakes

Reasonable starting points; tune on your own task and model.
SituationSuggested settingWhy
One correct answer (extraction, classification)Greedy, or temperature near 0Truncation does not matter if you always take the top token
General chat, support answersTop-p around 0.9 to 0.95, moderate temperatureNatural variety without tail nonsense
Creative writing, brainstormingTop-p around 0.95, temperature around 1Wide but still plausible choices
Need a strict upper bound on optionsTop-k (e.g. 40 to 50) plus top-pk caps the list when p would keep too many

Common mistakes Setting top-p = 1.0 and thinking you have filtered something (p = 1 keeps every token). Setting k = 1 and expecting variety (it is greedy decoding). Cranking temperature high and relying on top-p to clean up: the flatter distribution makes the nucleus huge. Tuning temperature, top-k and top-p all at once without testing; change one at a time. And remembering that some APIs expose only some of these parameters.

Real-world use Most LLM APIs and open-source inference servers expose temperature and top_p, and many also accept top_k. A support bot might run with low temperature and top-p 0.9 for steady answers, while a "suggest 5 taglines" feature uses higher temperature and top-p 0.95 to get genuinely different options.

Worked example, step by step

We have used each filter alone. In practice they run one after another, and each one changes the input of the next. Let us chain two of them on our hotel example: first top-k with k = 5, then top-p with p = 0.9.

Top-k = 5, then top-p = 0.9

  1. Top-k keeps five: Paris 0.50, Lyon 0.15, France 0.12, the 0.08, Nice 0.06. Their total is 0.91.
  2. Renormalize: Divide by 0.91: Paris 0.549, Lyon 0.165, France 0.132, the 0.088, Nice 0.066.
  3. Running total for top-p: On the new numbers: 0.549 → 0.714 → 0.846 → 0.934. The total passes 0.9 at the 4th token.
  4. Top-p keeps four: Paris, Lyon, France and "the". "Nice" is dropped, even though top-k had let it through.
  5. Renormalize again and sample: Divide by 0.934: Paris 0.588, Lyon 0.176, France 0.141, the 0.094.

Compare this with top-p = 0.9 alone, which kept five tokens in our earlier code. The chain keeps only four. Top-k removed some probability, renormalizing made every survivor a little bigger, and so the running total reached 0.9 one token sooner. Filters are not independent: the second one sees the output of the first.

Temperature sits even earlier in the chain, so it changes everything after it. Here is the size of the p = 0.9 nucleus for the same eight tokens when we apply a temperature first:

Finally, the lesson mentioned min-p. It keeps every token whose probability is at least a fraction of the top token's probability. With min-p = 0.1 and Paris at 0.50, the bar is 0.05. The table compares the three filters on the three distributions from our code:

Number of tokens kept (out of 8), computed from the lesson's three distributions
DistributionTop-k, k = 3Top-p, p = 0.9Min-p, 0.1
Hotel example (top token 0.50)356
Confident (top token 0.92)311
Flat (top token 0.16)378

Top-k never moves. Top-p and min-p both shrink when the model is sure and widen when it is not, but they measure "sure" differently: top-p looks at the total mass, min-p looks at the gap to the leader.

Practice: try it yourself

We give the long tail real size and measure the damage. Our vocabulary has 8 sensible tokens and 200 junk tokens that each have a tiny probability. We draw 10,000 tokens under six settings and count how many draws are junk.

practice_tail_filter.py

import numpy as np
rng = np.random.default_rng(3)
# 8 sensible tokens hold 94% of the probability; 200 junk tokens share the last 6%
head = np.array([0.47, 0.14, 0.11, 0.075, 0.056, 0.047, 0.028, 0.014])
tail = np.full(200, 0.0003)
probs = np.concatenate([head, tail])          # already sorted high -> low
is_junk = np.arange(len(probs)) >= len(head)
def truncate(p, k=None, top_p=None):
"""Keep the top k tokens, then the nucleus reaching top_p, then renormalize."""
p = p.copy()
if k is not None:
p[k:] = 0.0                           # works because p is sorted
p = p / p.sum()
if top_p is not None:
cum = np.cumsum(p)
n = np.searchsorted(cum, top_p) + 1   # smallest set reaching top_p
p[n:] = 0.0
p = p / p.sum()
return p
settings = [("pure sampling", {}), ("top-k = 50", {"k": 50}), ("top-k = 5", {"k": 5}),
("top-p = 0.9", {"top_p": 0.9}), ("top-p = 0.99", {"top_p": 0.99}),
("k = 50 then p = 0.9", {"k": 50, "top_p": 0.9})]
print("setting               kept   junk draws in 10,000")
for name, kwargs in settings:
p = truncate(probs, **kwargs)
draws = rng.choice(len(p), size=10_000, p=p)
print(f"{name:20s} {np.count_nonzero(p):5d}   {int(is_junk[draws].sum()):6d}")

Output:

setting               kept   junk draws in 10,000
pure sampling          208      625
top-k = 50              50      128
top-k = 5                5        0
top-p = 0.9              7        0
top-p = 0.99           175      501
k = 50 then p = 0.9      6        0

Now change it:

  • Spread the same 6% over more junk tokens: tail = np.full(2000, 0.00003). Predict the junk draws for "pure sampling" and for "top-k = 50". Which one changes a lot, and why?
  • Add a setting ("top-p = 0.95", {"top_p": 0.95}). The sensible tokens hold 94%. Predict whether junk gets in, and roughly how many tokens are kept.
  • Apply a temperature of 1.5 before truncating: after building probs, add probs = probs ** (1 / 1.5) and probs = probs / probs.sum(). Predict whether "top-p = 0.9" still gives zero junk draws.

Pause and think: Top-p = 0.9 alone kept 7 tokens. Adding top-k = 50 in front of it kept only 6. Top-k = 50 on its own keeps 50 tokens, so how can it make the result smaller?

Top-k = 50 removes 158 junk tokens, about 4.7% of the probability. Renormalizing divides every survivor by about 0.953, so each one grows a little. The first six sensible tokens sum to 0.898 before, just short of 0.9, and to about 0.943 after, which passes 0.9. The nucleus closes one token earlier. Every filter changes the numbers the next filter sees.

Pause and think: Top-p = 0.99 let 501 junk draws through, while top-p = 0.9 let none. The two settings look close. Why is the result so different here?

Because the sensible tokens hold only 94% of the probability. A nucleus of 0.9 fits inside that 94%, so it never touches the tail. A nucleus of 0.99 cannot be filled by sensible tokens alone: it must take another 5% from the tail, which means 167 junk tokens. What matters is not how close p is to 1, but whether p is above or below the share held by the plausible tokens at that step.

Key takeaways

  • Greedy decoding is dull and repetitive; sampling from every token lets in tail nonsense.
  • Top-k keeps a fixed number of the most likely tokens, then renormalizes.
  • Top-p keeps the smallest set of top tokens whose probabilities sum to at least p, so it adapts to confidence.
  • Temperature reshapes the distribution before truncation; high temperature makes the top-p nucleus bigger.
  • A common default for open-ended text is top-p around 0.9 to 0.95 with moderate temperature; use greedy for single-answer tasks.

Key terms

  • Greedy decoding: Always choosing the single most likely next token.
  • Top-k sampling: Sampling only from the k most likely tokens after renormalizing.
  • Top-p (nucleus) sampling: Sampling from the smallest set of top tokens whose probabilities sum to at least p.
  • Long tail: The many low-probability tokens that together can hold noticeable probability mass.
  • Renormalize: Rescale the kept probabilities so they sum to 1 again.
  • Cumulative probability: The running total of probabilities when tokens are sorted from most to least likely.

← 5.1 Temperature Sampling: Dialing Up or Down Creativity · 5.3 Token Streaming: Rendering Outputs as They Arrive →