Lesson 13.3 · 26 min
Prefill-Decode Disaggregation: Splitting the Two Phases
What if the GPU that reads your prompt and the GPU that writes your answer were two different machines, and that made both faster?
In short: Prefill is compute-heavy and decode is memory-heavy, so when both share one GPU they get in each other's way: a long prompt arriving can freeze everyone else's token stream. Prefill-decode disaggregation runs the two phases on separate GPU pools and ships the KV cache between them. It gives steadier TTFT and TPOT and lets each pool be tuned on its own, at the cost of KV transfer, a fast network and more operational complexity.
How an LLM answers a request, and what the KV cache is
A request to an LLM goes through two phases. Prefill reads the entire prompt in one parallel forward pass. Decode then writes the answer one token per forward pass, each new token depending on the one before it.
The KV cache connects them. In every attention layer, each token produces a key vector (what it offers for matching) and a value vector (what it contributes when attended to). Prefill computes keys and values for all prompt tokens and stores them. Each decode step computes them only for the newest token, appends them, and attends over the whole stored cache. Without the cache, every decode step would redo all the prompt work.
Think of it like a restaurant kitchen Prep cooks chop vegetables in big batches; line cooks plate dishes one at a time, fast and steadily. If one person did both, every time a huge prep job arrived the plates would stop going out. Big kitchens separate the two stations and pass the prepped ingredients across. The prepped ingredients are the KV cache.
Our running example is a support chatbot where some customers paste 20,000-token log files while dozens of others are mid-way through receiving short answers. We care about both how fast the first token appears and how smoothly the rest flows.
Prefill is compute-heavy, decode is memory-heavy
A GPU forward pass must read the model's weights from memory and then do math with them. In prefill, one weight read is shared by thousands of prompt tokens, so the math dominates: prefill is compute-bound. In decode, each request contributes just one token per pass, so the GPU spends most of its time reading weights and the KV cache: decode is memory-bandwidth-bound.
- Prefill wants many FLOPs: it benefits from more compute and splitting work across GPUs (tensor parallelism).
- Decode wants high memory bandwidth and lots of memory capacity: big batches of users share each weight read, and each user's KV cache must fit.
- A prefill of a long prompt can take tens or hundreds of milliseconds; a decode step typically takes a few tens of milliseconds.
Same model, two workloads Nothing about the weights is different between phases. Only the shape of the work differs: matrix × matrix for prefill, many matrix × vector products for decode.
The problem when both run on the same GPU
In a classic co-located server, one GPU (or one group of GPUs) runs both phases. A scheduler forms each forward pass from whatever work is waiting. When a new long prompt arrives, there are two choices, and both hurt someone:
- Run the prefill now. Everyone currently decoding waits for it. Their next token arrives late, so their stream stutters.
- Keep decoding first. The new user's prefill waits in the queue, so their first token arrives late.
This is called prefill-decode interference. Let us simulate it. One user is decoding at 20 ms per step. Long prompts arrive every 100 ms and each prefill takes 60 ms (illustrative numbers). A co-located GPU runs any waiting prefill before the next decode step. A disaggregated setup sends prefills to another GPU.
interference.py
# Toy timeline: one user is decoding while new prompts keep arriving.
# Times in ms, illustrative: prefill of a long prompt = 60 ms,
# one decode step = 20 ms, KV transfer between GPUs = 5 ms.
PREFILL, DECODE, TRANSFER = 60, 20, 5
arrivals = [0, 100, 200, 300] # new requests (long prompts)
def colocated(n_tokens=20):
# One GPU. A waiting prefill runs before the next decode step.
t, gaps, queue = 0, [], list(arrivals[1:])
for _ in range(n_tokens):
start = t
while queue and queue[0] <= t: # prefill jumps in first
queue.pop(0); t += PREFILL
t += DECODE
gaps.append(t - start) # time between our two tokens
return gaps
def disaggregated(n_tokens=20):
# Prefills go to a separate GPU; decode GPU only decodes.
# The KV transfer happens once, before our first token.
return [DECODE] * n_tokens, TRANSFER
co = colocated()
dis, xfer = disaggregated()
print("co-located gaps (ms): ", co)
print("disaggregated gaps (ms):", dis)
print(f"co-located avg TPOT {sum(co)/len(co):.0f} ms, worst {max(co)} ms")
print(f"disaggregated avg TPOT {sum(dis)/len(dis):.0f} ms, worst {max(dis)} ms"
f" (+{xfer} ms one-time KV transfer on TTFT)")Output:
co-located gaps (ms): [20, 20, 20, 20, 20, 80, 20, 80, 20, 80, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20] disaggregated gaps (ms): [20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20, 20] co-located avg TPOT 29 ms, worst 80 ms disaggregated avg TPOT 20 ms, worst 20 ms (+5 ms one-time KV transfer on TTFT)
Three of our user's tokens took 80 ms instead of 20 ms: a 4× stall each time someone else's prompt arrived. The average looks tolerable (29 ms), but users notice the worst case, and latency targets are usually set on high percentiles such as p99.
TTFT vs TPOT: two promises that pull apart
TTFT (time to first token) is how long a user waits before anything appears; it is mostly queueing plus prefill. TPOT (time per output token) is the gap between later tokens; it is set by decode step time and by any interruptions. A chat product usually promises both, for example TTFT under 1 s and TPOT under 50 ms.
On a co-located GPU these targets fight each other. Prioritising prefill improves TTFT but causes TPOT spikes. Prioritising decode smooths TPOT but makes new users wait. To satisfy both, operators often over-provision GPUs and run them below capacity. The DistServe paper (OSDI 2024) framed this in terms of goodput: requests per second served while meeting both latency targets, and argued that interference wastes much of a co-located cluster's goodput.
Pause and think: If we simply run fewer requests per GPU, does interference go away?
It shrinks, because fewer prompts arrive to interrupt each decode batch, but it does not disappear and we pay with low utilisation and more GPUs. That is the over-provisioning trap.
The naive approaches and their issues
| Approach | What it does | Issue |
|---|---|---|
| Prefill-first scheduling | Always run waiting prefills immediately | Good TTFT, but decode streams stall (our 80 ms spikes) |
| Decode-first scheduling | Finish decode steps before starting prefills | Smooth TPOT, but new users can wait a long time for the first token |
| Chunked prefill | Split a long prompt into chunks (e.g. 512 tokens) and mix one chunk with each decode batch | Much smaller spikes, but each step still carries prefill work; long prompts take more steps to finish; both phases still share one hardware configuration |
| Over-provisioning | Keep load low so collisions are rare | Expensive: GPUs sit underused |
Chunked prefill deserves credit: it is widely used in engines like vLLM and SGLang, and for many deployments it is enough. But it only softens the interference. The two phases still compete for the same GPUs, and we cannot pick different parallelism or batch sizes for each phase.
What prefill-decode disaggregation is and how it works
Prefill-decode disaggregation (often “PD disaggregation”) splits the serving fleet into two pools: prefill workers that only run prefills, and decode workers that only run decode steps. After a prefill finishes, the request's KV cache is sent over a fast interconnect to a decode worker, which continues the generation.
Walkthrough: a customer pastes a 20,000-token log
- Arrive and route: The router sends the request to the least-busy prefill worker.
- Prefill: The prefill worker processes all 20,000 tokens in one pass (possibly split across several GPUs with tensor parallelism) and samples the first token.
- Ship the cache: The KV cache for 20,000 tokens is transferred to a decode worker. With GQA (8 KV heads), 32 layers, head dim 128 and FP16, that is 128 KiB per token, about 2.4 GiB in total, so a fast link matters.
- Join a decode batch: The decode worker slots the request into its continuous batch alongside dozens of other conversations.
- Generate smoothly: Because no prefill ever runs on this worker, every user's inter-token gap stays close to one decode step.
Tune each pool for its job Prefill workers can use more tensor parallelism to cut TTFT; decode workers can use large batches and lots of memory for KV caches. Teams can also change the ratio of prefill to decode GPUs as the traffic mix changes, for example more prefill workers for a summarisation-heavy app with long inputs.
Advantages, disadvantages and where it fits
Real-world use Research systems DistServe and Splitwise (both 2024) showed the benefits of separating phases, and Moonshot AI described Mooncake, a KV-cache-centric disaggregated architecture serving its Kimi chatbot. Open-source frameworks have added support: NVIDIA Dynamo is built around disaggregated serving, and vLLM and SGLang offer PD-disaggregation modes. Exact features change quickly, so check current docs.
Where it is overkill If you serve a small model on one or two GPUs, prompts are short, or traffic is light, disaggregation adds a transfer hop and moving parts for little gain. Short prompts mean short prefills, so there is little interference to remove. Measure p99 TPOT spikes first; if chunked prefill already meets your targets, stay co-located.
Pause and think: Our chatbot's prompts get much shorter after we add retrieval that sends only the relevant 1,000 tokens. Does the case for disaggregation get stronger or weaker?
Weaker. Shorter prompts mean shorter prefills that interrupt decode less, and a smaller KV cache to transfer, so the benefit shrinks relative to the added complexity.
Going one level deeper
The comparison above says the pool ratio must track the traffic mix. Let us see what that means with small numbers. All figures here are illustrative. Our chatbot receives 10 requests per second. Each prefill keeps one prefill worker busy for 0.25 s. Each answer is 200 tokens at 20 ms per step, so a request spends 4 s in decode. One decode worker can hold 32 requests in its batch.
How many workers of each kind?
- Prefill work per second: 10 requests × 0.25 s = 2.5 seconds of prefill work arrive every second. One worker can do 1 second of work per second, so we need 3 prefill workers (2.5 rounded up).
- Requests in decode at once: 10 requests arrive per second and each stays 4 s, so about 10 × 4 = 40 requests are decoding at any moment.
- Decode workers: 40 requests ÷ 32 per worker = 1.25, so we need 2 decode workers.
- The ratio: 3 prefill workers to 2 decode workers. This ratio comes from the traffic, not from the model.
- Change the traffic: If prompts double in length, prefill work doubles to 5 seconds per second: 5 prefill workers, decode unchanged. If answers double in length, 80 requests decode at once: 3 decode workers, prefill unchanged.
| Traffic | Prefill workers | Decode workers |
|---|---|---|
| Baseline: 10 requests/s | 3 | 2 |
| Prompts twice as long | 5 | 2 |
| Answers twice as long | 3 | 3 |
| Twice as many requests | 5 | 3 |
Now the failure cases. If the prefill pool is too small, requests queue before prefill: TTFT climbs while TPOT stays perfectly smooth. If the decode pool is too small, finished prefills wait for a free slot, or caches no longer fit in memory: prefill workers look idle, yet first tokens are still late. If the link between pools is slow, every TTFT grows by the same extra amount, and it grows with prompt length because a longer prompt has a bigger cache to copy.
So when a disaggregated cluster misses its targets, we first ask which number is bad and where the requests are waiting. Adding workers to the wrong pool changes nothing.
Practice: try it yourself
We will simulate the prefill pool on its own. Long prompts arrive at a steady pace, each one waits for the next free prefill worker, runs its prefill, and then pays a short KV transfer. We print every request's TTFT for one, two and three prefill workers. The timings are illustrative.
practice_prefill_pool.py
# Sizing the prefill pool: how many prefill workers keep TTFT steady?
# Times in ms, illustrative.
import heapq
PREFILL, TRANSFER = 250, 10 # one long prompt; KV copy to a decode worker
ARRIVAL_GAP, N_REQUESTS = 100, 12 # a new request arrives every 100 ms
def ttfts(n_workers):
free_at = [0] * n_workers # time at which each prefill worker is free
heapq.heapify(free_at)
out = []
for i in range(N_REQUESTS):
arrive = i * ARRIVAL_GAP
start = max(arrive, heapq.heappop(free_at)) # wait for the next free worker
done = start + PREFILL
heapq.heappush(free_at, done)
out.append(done + TRANSFER - arrive) # queue + prefill + transfer
return out
for workers in (1, 2, 3):
t = ttfts(workers)
load = PREFILL / (ARRIVAL_GAP * workers) # work arriving / capacity
print(f"{workers} prefill worker(s): load={load:.2f} "
f"TTFT first={t[0]} ms, last={t[-1]} ms")
print(" all TTFTs:", t)Output:
1 prefill worker(s): load=2.50 TTFT first=260 ms, last=1910 ms all TTFTs: [260, 410, 560, 710, 860, 1010, 1160, 1310, 1460, 1610, 1760, 1910] 2 prefill worker(s): load=1.25 TTFT first=260 ms, last=510 ms all TTFTs: [260, 260, 310, 310, 360, 360, 410, 410, 460, 460, 510, 510] 3 prefill worker(s): load=0.83 TTFT first=260 ms, last=260 ms all TTFTs: [260, 260, 260, 260, 260, 260, 260, 260, 260, 260, 260, 260]
With one or two workers the load is above 1.0 and TTFT rises with every new request. With three workers every request gets 260 ms: the prefill plus the transfer and no queueing. Now change it:
- Set
PREFILL = 190. Predict the load for two workers, and whether two workers are now enough, before you run it. - Set
TRANSFER = 200to model a slow link. Predict what happens to the three-worker TTFTs. Does adding a fourth worker help? - Set
N_REQUESTS = 40with the original timings. Predict the last TTFT for two workers. Does the queue settle at some level or keep growing?
Pause and think: With two workers the load is 1.25. Why does TTFT keep rising instead of settling at a higher but steady value?
A load above 1.0 means work arrives faster than the pool can finish it. Every new request finds a slightly longer queue than the one before, so the wait grows without limit. Here each request adds 25 ms: the pool finishes one prefill every 125 ms on average, but one arrives every 100 ms.
Pause and think: After moving to two pools, our TPOT is smooth, but during a burst of long pasted logs the first token arrives late. A teammate wants to add decode workers. Is that the right fix?
No. Smooth TPOT says the decode pool is coping. Late first tokens during a burst of long prompts point to queueing in the prefill pool, or to slow KV transfers. We should add prefill workers or check the link, not add decode workers.
Key takeaways
- Prefill is compute-bound and decode is memory-bound; on shared GPUs they interfere.
- Interference shows up as TTFT delays or TPOT spikes, and latency targets care about the worst cases.
- Disaggregation runs prefill and decode on separate pools and transfers the KV cache between them.
- It enables independent tuning and scaling of each phase, at the cost of KV transfer and complexity.
- For small deployments or short prompts, co-located serving with chunked prefill is usually enough.
Key terms
- Co-located serving: Running both prefill and decode on the same GPUs.
- Prefill-decode disaggregation: Serving prefill and decode on separate GPU pools, moving the KV cache from one to the other.
- Interference: Slowdown one phase causes the other when they share a GPU, such as a prefill stalling decode steps.
- Chunked prefill: Splitting a long prompt's prefill into smaller pieces that are mixed into decode batches.
- Goodput: Throughput counting only requests that meet their latency targets.
- KV transfer: Copying a request's key/value cache from a prefill worker to a decode worker over a fast interconnect.
← 13.2 Prefill vs Decode: Two Distinct Phases of LLM Inference · 13.4 The KV Cache: Avoiding Redundant Attention Computation →