Modern AI Engineering

Lesson 13.2 · 27 min

Prefill vs Decode: Two Distinct Phases of LLM Inference

Why can a GPU read your 4,000-token prompt in a fraction of a second, yet take several seconds to write a 300-token answer?

In short: Every LLM request runs in two phases. Prefill processes the whole prompt in one parallel pass and builds the KV cache; it is limited by raw compute. Decode then produces one token per step, reading all the weights and the cache each time; it is limited by memory bandwidth. Because the phases stress different hardware limits, they have different metrics (TTFT vs TPOT) and different optimizations.

What LLM inference is

Inference is running a trained model to get an answer. For an LLM, a request arrives with a prompt (the system instructions, chat history and user message, all turned into tokens), and the model returns generated tokens until it emits an end-of-sequence token or hits a length limit.

Our running example is a support chatbot. A customer pastes a 2,000-token error log and asks “Why did my upload fail?”. The bot answers in about 200 tokens. Behind that single exchange there are two very different kinds of work, and serving systems treat them separately.

Think of it like a student in an exam First the student reads the whole question paper. Reading is fast because the eyes take in many words at once. Then they write the answer one word at a time, and before each word they glance at their notes. Reading the paper is prefill; writing word by word is decode; the notes are the KV cache.

The two phases: prefill and decode

Prefill (also called the prompt phase) takes all prompt tokens at once and runs them through every layer in a single forward pass. Because the model already knows every prompt token, it can compute all of them in parallel, using a causal mask so each token only attends to earlier ones. Prefill produces two things: the KV cache for every prompt token, and the probabilities for the very first output token.

Decode (also called the generation phase) produces the rest of the answer. Each decode step feeds in just one new token, the one picked in the previous step, computes its query, key and value in every layer, appends the key and value to the cache, attends over the whole cache and predicts the next token. Decode cannot be parallelised across output tokens within one request, because token 5 depends on what token 4 turned out to be.

The KV cache is the bridge between the two phases. Prefill writes it, decode reads it at every step and appends to it. Without the cache, every decode step would have to redo the prefill work for the entire text so far.

A few decode steps, slowly

Let us shrink the example. The prompt is the 4 tokens “Why did upload fail”. Prefill processes all 4 and the cache now holds 4 rows per layer.

Decode, step by step

  1. After prefill: Cache: 4 rows. The last prompt position predicts output token 1, say “The”. Time so far is the TTFT.
  2. Decode step 1: Input: “The”. Compute its Q, K, V in each layer; append K and V (cache: 5 rows). Its query attends over 5 keys. Output: “file”.
  3. Decode step 2: Input: “file”. Append (cache: 6 rows). Attend over 6 keys. Output: “was”.
  4. Decode step 3: Input: “was”. Append (cache: 7 rows). Attend over 7 keys. Output: “too”.
  5. And so on: Each step does a small amount of math for one token but must read all model weights plus a cache that keeps growing.

Pause and think: Pause and predict: in decode step 50 of our real example (2,000-token prompt), how many cached rows does each layer's attention read?

About 2,050: the 2,000 prompt rows plus the 50 generated tokens (the newest one included). The cache read grows by one row every step.

Prefill vs decode side by side

Why the split matters: compute-bound vs memory-bound

A GPU has two key limits. Compute throughput is how many floating-point operations (FLOPs) per second it can do. Memory bandwidth is how many bytes per second it can move from its main memory (HBM) to its compute units. An NVIDIA H100 SXM is rated around 989 teraFLOPs of dense BF16 math and about 3.35 terabytes per second of bandwidth.

Arithmetic intensity is the number of FLOPs we do per byte we move. If it is high, the math units are the bottleneck: the work is compute-bound. If it is low, the math units sit idle waiting for data: the work is memory-bound. The crossover, called the ridge point, is peak FLOPs divided by bandwidth: about 989e12 / 3.35e12 ≈ 295 FLOPs per byte for the H100.

So intensity is roughly the number of tokens processed per pass. A prefill of 2,000 tokens has intensity near 2,000, far above 295: compute-bound. A decode step for one user has intensity near 1: badly memory-bound. The GPU spends almost all of that step just streaming 16 GB of weights through its memory system.

roofline.py

# Is a step compute-bound or memory-bound? A back-of-envelope roofline.
params = 8e9                 # 8B-parameter model, FP16 weights
weight_bytes = params * 2    # 16 GB read from GPU memory per forward pass
peak_flops = 989e12          # approx. H100 SXM dense BF16 peak
mem_bw = 3.35e12             # approx. H100 SXM HBM bandwidth, bytes/s
ridge = peak_flops / mem_bw  # FLOPs per byte needed to be compute-bound
def step_time(tokens):
flops = 2 * params * tokens            # ~2 FLOPs per weight per token
t_compute = flops / peak_flops
t_memory = weight_bytes / mem_bw       # weights are read once per pass
intensity = flops / weight_bytes       # FLOPs per byte moved
bound = "compute" if t_compute > t_memory else "memory"
return max(t_compute, t_memory), intensity, bound
print(f"ridge point: {ridge:.0f} FLOPs/byte")
for label, n in [("decode, 1 user", 1), ("decode, 32 users", 32),
("prefill, 512-token prompt", 512),
("prefill, 4096-token prompt", 4096)]:
t, ai, bound = step_time(n)
print(f"{label:27s} tokens={n:5d} intensity={ai:6.0f} "
f"time={t*1e3:7.2f} ms  {bound}-bound  "
f"({t*1e3/n:.3f} ms/token)")

Output:

ridge point: 295 FLOPs/byte
decode, 1 user              tokens=    1 intensity=     1 time=   4.78 ms  memory-bound  (4.776 ms/token)
decode, 32 users            tokens=   32 intensity=    32 time=   4.78 ms  memory-bound  (0.149 ms/token)
prefill, 512-token prompt   tokens=  512 intensity=   512 time=   8.28 ms  compute-bound  (0.016 ms/token)
prefill, 4096-token prompt  tokens= 4096 intensity=  4096 time=  66.26 ms  compute-bound  (0.016 ms/token)

Look at the decode lines. One user and 32 users take the same 4.78 ms per step, because the step is dominated by reading the weights, and the weights are read once no matter how many users share the pass. That is why batching is the main tool for decode throughput. Prefill, by contrast, costs about 0.016 ms per token regardless of prompt size: it is already saturating the math units.

The KV cache adds memory traffic too Our simple model only counts weight reads. In real decode, each step also reads the request's entire KV cache. With long contexts and big batches, cache reads can rival or exceed weight reads, which is why KV cache compression and GQA speed up decode.

The key metrics: TTFT, TPOT, throughput, latency

MetricWhat it measuresMostly set byOur chatbot target (illustrative)
TTFT (time to first token)From request arrival to the first output tokenQueueing + prefillUnder 1 s
TPOT (time per output token), also ITL (inter-token latency)Average gap between consecutive output tokensDecode step timeUnder 50 ms (faster than reading speed)
End-to-end latencyArrival to last tokenTTFT + TPOT × (output tokens − 1)Under 12 s for 200 tokens
ThroughputTokens (or requests) per second across all usersBatch size and GPU utilisationAs high as possible within the latency targets

There is a built-in tension. Bigger decode batches raise throughput (more tokens per weight read) but each step gets a bit slower, so TPOT rises. Long prefills in the same batch can delay everyone's next token. Serving is about meeting latency targets, often called SLOs (service-level objectives), while keeping throughput high.

Pause and think: Users complain the bot “takes ages to start answering” on long pasted logs, but once it starts, text flows smoothly. Which phase and metric should we look at?

Prefill and TTFT. A long prompt means a long prefill (plus any queueing) before the first token. Decode, and therefore TPOT, is fine.

Optimization techniques mapped to each phase

Real-world use Open-source engines such as vLLM, SGLang and TensorRT-LLM implement most of these: continuous batching, paged KV memory, chunked prefill, prefix caching and speculative decoding. API providers expose the same split in pricing: input (prompt) tokens are usually much cheaper per token than output tokens, reflecting that prefill is far more efficient per token than decode.

Worked example, step by step

Let us add up the time for one whole request by hand, using the ideal numbers from roofline.py: about 0.016 ms per prompt token in prefill, and 4.78 ms per decode step. Our customer sends a 2,000-token log and gets a 200-token answer. These are lower bounds from a simple model, not measurements, but the proportions are what we care about.

Where the time goes for one chat request

  1. Prefill: 2,000 tokens × 0.016 ms ≈ 32 ms. All 2,000 tokens share one read of the weights, so this pass is compute-bound.
  2. First token: Prefill also gives us output token 1. With no queueing, TTFT ≈ 32 ms.
  3. Decode: The other 199 tokens need 199 steps. 199 × 4.78 ms ≈ 951 ms. Every step reads all 16 GB of weights to produce one token.
  4. Add them up: End-to-end ≈ 32 + 951 = 983 ms. Decode handled about 9% of the tokens (200 of 2,200) but took about 97% of the time.
  5. Check the formula: TTFT + TPOT × (N − 1) = 32 + 4.78 × 199 ≈ 983 ms. The hand sum and the formula agree.
Ideal time split for three workloads on the lesson's roofline model (lower bounds, not measurements)
WorkloadPrompt → output tokensPrefillDecodeDecode share of time
Chat2,000 → 200≈ 32 ms≈ 951 ms≈ 97%
Summariser20,000 → 100≈ 324 ms≈ 473 ms≈ 59%
Story writer200 → 1,000≈ 4.8 ms≈ 4,775 ms≈ 99.9%

Look closely at the story writer's prefill. 200 tokens × 0.016 ms would be 3.2 ms, but the table says 4.8 ms. A pass can never be faster than one full read of the weights, which takes 4.78 ms. With only 200 tokens the intensity is about 200, below the ridge point of 295, so even this prefill is memory-bound. Prefill is compute-bound only when the prompt is long enough.

A common mistake is to judge a workload by its total token count. The chat and a story of 200 → 2,000 tokens both move 2,200 tokens, yet the second takes about ten times longer, because every output token pays for its own pass.

Practice: try it yourself

We will build a small latency calculator. It takes a prompt length, an output length and a decode batch size, and returns TTFT, TPOT, total time and server throughput. The timing constants are illustrative. The point is to see how the two phases and the batch size trade against each other.

practice_latency_budget.py

# One request's latency from its two phases, and how batch size shifts it.
# All timings are illustrative, not measured.
PREFILL_MS_PER_TOKEN = 0.02     # prefill is compute-bound: cost per prompt token
DECODE_BASE_MS = 5.0            # reading the weights once per decode step
DECODE_PER_USER_MS = 0.05       # small extra work per request in the batch
def request_metrics(prompt_tokens, output_tokens, batch, queue_ms=0.0):
ttft = queue_ms + prompt_tokens * PREFILL_MS_PER_TOKEN
tpot = DECODE_BASE_MS + DECODE_PER_USER_MS * batch
total = ttft + tpot * (output_tokens - 1)
throughput = batch * 1000 / tpot           # tokens/s across all users
return ttft, tpot, total, throughput
workloads = [("chat", 2000, 200), ("summarise", 20000, 100), ("story", 200, 1000)]
print("workload   batch  TTFT ms  TPOT ms  total s  decode share  server tok/s")
for name, p, n in workloads:
for batch in (1, 32):
ttft, tpot, total, thr = request_metrics(p, n, batch)
share = 1 - ttft / total               # part of the wait spent decoding
print(f"{name:10s} {batch:5d} {ttft:8.0f} {tpot:8.2f} {total/1000:8.2f} "
f"{share:12.0%} {thr:13.0f}")

Output:

workload   batch  TTFT ms  TPOT ms  total s  decode share  server tok/s
chat           1       40     5.05     1.04          96%           198
chat          32       40     6.60     1.35          97%          4848
summarise      1      400     5.05     0.90          56%           198
summarise     32      400     6.60     1.05          62%          4848
story          1        4     5.05     5.05         100%           198
story         32        4     6.60     6.60         100%          4848

Going from batch 1 to batch 32 made every user's TPOT a little worse (5.05 ms to 6.60 ms) while the server's output rose about 24 times. Now change it:

  • Pass queue_ms=500 for the chat workload. Predict first: which of the four numbers change, and which stay exactly the same?
  • Set DECODE_PER_USER_MS = 0.5. Predict whether batch 32 still gives more server tokens per second than batch 1, and how much TPOT each user now sees.
  • Add a workload ("swap", 200, 2000) next to the chat one. Both move 2,200 tokens. Predict which is slower and by roughly what factor before you run it.

Pause and think: In the practice output, the summariser has the highest TTFT but the lowest total time. How can both be true?

TTFT depends on the prompt, and its prompt is the longest (20,000 tokens). Total time is mostly decode steps, and it writes the fewest output tokens (100). A long wait to start and a short wait to finish are set by different phases.

Pause and think: We keep raising the decode batch size because server tokens per second keeps going up. What tells us to stop?

Each user's TPOT. Every extra request in the batch makes the step slightly longer, so every user's tokens arrive a little more slowly. We stop when TPOT reaches our latency target, or when the KV caches no longer fit in memory, whichever comes first.

Conclusion and common mistakes

Every request is one prefill followed by many decode steps. Prefill is a big parallel pass that fills the KV cache and sets TTFT. Decode is a long chain of tiny passes that read the weights and cache over and over and sets TPOT. Knowing which phase dominates your workload tells you which optimizations will pay off: a summariser with 20k-token inputs and short outputs is prefill-heavy, while a creative-writing assistant with short prompts and long answers is decode-heavy.

Common mistakes Reporting only “tokens per second” without saying whether it is per user or across the whole server. Assuming a faster GPU in FLOPs will speed up decode (bandwidth matters more). And benchmarking with short prompts only, which hides prefill and TTFT problems that real users with long documents will hit.

Key takeaways

  • Each request = one parallel prefill pass over the prompt + one decode step per output token.
  • The KV cache is written in prefill and read (and extended) at every decode step.
  • Prefill is compute-bound; decode is memory-bandwidth-bound because it reads all weights for few tokens.
  • TTFT reflects queueing + prefill; TPOT reflects decode; end-to-end ≈ TTFT + TPOT × (N − 1).
  • Optimize prefill by cutting FLOPs and decode by cutting bytes per token or batching more users.

Key terms

  • Prefill: The phase that processes all prompt tokens in one parallel pass, building the KV cache and producing the first token.
  • Decode: The phase that generates output tokens one per forward pass, reusing and extending the KV cache.
  • Arithmetic intensity: FLOPs performed per byte moved from memory; low values mean a memory-bound workload.
  • TTFT: Time to first token: delay from request arrival to the first output token.
  • TPOT: Time per output token: average gap between consecutive generated tokens.
  • Throughput: Total tokens or requests a server produces per second across all users.
  • SLO: Service-level objective: a target such as “TTFT under 1 s for 99% of requests”.

← 13.1 LLM Inference Optimization: The Full Landscape · 13.3 Prefill-Decode Disaggregation: Splitting the Two Phases →