Lesson 9.3 · 22 min
Prompt Caching: Reusing Computation Across API Calls
If our chatbot sends the same 20,000-token manual with every question, why should the model read it from scratch every single time?
In short: When an LLM reads a prompt it computes internal key and value tensors for every token. Prompt caching stores those tensors for a prompt prefix so the next request that starts with exactly the same tokens can skip that work. The result is lower latency and much cheaper input tokens, as long as we put the stable content first and keep it byte-for-byte identical.
What is a prompt, and how does an LLM read it?
A prompt is everything we send to a large language model (LLM) in one request: the system prompt (standing instructions), tool definitions, documents, earlier conversation turns and the new user message. The model turns this text into tokens, small pieces of text, and processes them in two phases.
- Prefill. The model reads all prompt tokens. In every layer, for every token, it computes a key vector and a value vector (the K and V of attention). These are stored in the KV cache, a block of GPU memory for this request.
- Decode. The model generates the answer one token at a time. Each new token attends to all the stored keys and values instead of recomputing them.
Prefill work grows with prompt length. A 24,000-token prompt means 24,000 tokens' worth of keys and values in every layer before the first answer token can appear. That is why long prompts have a slow time to first token (TTFT) and why providers charge for input tokens.
Think of it like a chef's mise en place A restaurant does not chop onions from scratch for every order. It preps the common ingredients once, keeps them ready for a while, and only cooks the part that is unique to each dish. Prompt caching preps the shared start of our prompts once and reuses it for every order that begins the same way.
What is prompt caching, and why do we need it?
Prompt caching (also called prefix caching or context caching) means the provider keeps the KV cache of a prompt's beginning after a request finishes. When a later request starts with the same prefix, the model loads those stored keys and values and only runs prefill on the new tokens at the end.
Our running example: Acme's support bot. Every request contains a 4,000-token system prompt plus a 20,000-token product manual, followed by a 50-token customer question. Without caching, every question pays to process 24,050 tokens. With caching, after the first request only the 50 new tokens need full processing.
- Chatbots resend the whole conversation on every turn; each turn shares everything before it with the previous turn.
- RAG and document Q&A resend the same long document for many questions.
- Agents resend a long system prompt and tool list on every step of their loop, often dozens of times per task.
- Coding assistants resend large chunks of the same codebase again and again.
In all these cases most of each prompt is a repeat. Caching turns that repeated work into a cheap lookup.
The core idea: a prefix's keys and values never change
Why is reuse even allowed? Because LLMs use causal attention: token number i can only look at tokens 1 to i, never at later tokens. So the keys and values for the first 24,000 tokens depend only on those 24,000 tokens. Whatever we append afterwards cannot change them. Two prompts that share the same first N tokens have identical KV tensors for those N tokens.
What happens on a cached request
- Hash the prefix: The serving system computes a fingerprint (hash) of the prompt's leading tokens, often in fixed-size blocks.
- Look it up: If a stored KV cache with that fingerprint exists and has not expired, it is a cache hit. Otherwise it is a cache miss.
- Reuse or compute: On a hit, the stored keys and values are loaded. On a miss, prefill runs normally and the result is saved for next time (a cache write).
- Prefill only the new part: The model runs prefill on the tokens after the cached prefix, here the 50-token question.
- Decode as usual: Answer generation is unchanged. Caching changes speed and cost, never the content of the answer.
The exact-prefix rule
A cache hit needs the prompt to match exactly, from the very first token, up to the cached point. Matching is on tokens, not meaning. One changed character changes the tokens there, and because every later key and value depends on all earlier tokens, everything after that point must be recomputed.
- Changing “Acme” to “ACME” in the system prompt → miss for the whole prompt.
- Putting today's date or a request ID at the top of the prompt → a miss on every new date or ID.
- Reordering tool definitions or JSON keys between requests → miss.
- Changing only the user question at the end → hit on everything before it.
Order the prompt from most stable to least stable Tools and system instructions first, then large reference documents and few-shot examples, then the conversation history, and the new user message last. Anything that changes per request (timestamps, user names, retrieved snippets) goes as late as possible.
Pause and think: Pause and predict: we add “Current time: 10:42:07” as the first line of the system prompt. What happens to our cache hit rate?
It drops to roughly zero. The first tokens differ on every request, so no request shares a prefix with an earlier one. Moving the timestamp to the end (or removing it) restores the hits.
Cache write vs cache read, and TTL
Providers price the two events differently. A cache write happens on a miss: the prefix is processed and stored. A cache read happens on a hit: the stored prefix is reused. Stored caches do not live forever. The TTL (time to live) is how long a cache entry survives without being used; each hit usually resets the clock.
| Provider | How it is turned on | Pricing idea | Lifetime |
|---|---|---|---|
| Anthropic (Claude) | Explicit: mark cache breakpoints with cache_control | Writes cost more than normal input (1.25× for the 5-minute TTL); reads cost about 0.1× | 5 minutes by default, refreshed on each hit; a 1-hour option at a higher write price |
| OpenAI | Automatic for long prompts (from 1,024 tokens) | No write surcharge; cached input tokens are discounted, by an amount that depends on the model | Typically minutes of inactivity; longer at off-peak times |
| Google (Gemini) | Implicit (automatic) caching on recent models, plus explicit context caches | Discounted cached tokens; explicit caches also pay for storage time | Explicit caches have a TTL we set |
| Self-hosted (vLLM, SGLang) | Automatic prefix caching in the server | No price, just saved GPU work | Until evicted from GPU memory |
There is usually a minimum length to cache (for example around 1,024 tokens on many models), because caching a tiny prefix saves almost nothing.
Code you can run: a toy prefix cache with prices
This simulation hashes the exact prefix text as the cache key and applies illustrative prices: $3 per million input tokens, writes at 1.25×, reads at 0.10×.
prefix_cache.py
import hashlib
BASE = 3.00 / 1_000_000 # illustrative $ per input token
WRITE, READ = 1.25, 0.10 # illustrative multipliers: cache write, cache read
cache = {} # prefix hash -> tokens stored
def run(prefix_blocks, new_tokens):
# The cache key is a hash of the EXACT prefix text
text = "".join(t for t, _ in prefix_blocks)
n_prefix = sum(n for _, n in prefix_blocks)
key = hashlib.sha256(text.encode()).hexdigest()[:10]
if key in cache:
status, mult = "HIT ", READ
else:
cache[key] = n_prefix
status, mult = "MISS", WRITE
cost = n_prefix * BASE * mult + new_tokens * BASE
plain = (n_prefix + new_tokens) * BASE
print(f"{status} key={key} cost=${cost:.4f} (no caching: ${plain:.4f})")
return cost
system = ("You are the support bot for Acme...", 4_000)
manual = ("<product manual text>", 20_000)
total = sum(run([system, manual], new_tokens=50) for _ in range(4))
print(f"4 calls: ${total:.4f} with caching vs ${4 * 24_050 * BASE:.4f} without")
# Change ONE character in the system prompt -> new hash -> cache miss
run([("You are the support bot for ACME...", 4_000), manual], new_tokens=50)Output:
MISS key=6a42db1fb8 cost=$0.0902 (no caching: $0.0722) HIT key=6a42db1fb8 cost=$0.0074 (no caching: $0.0722) HIT key=6a42db1fb8 cost=$0.0074 (no caching: $0.0722) HIT key=6a42db1fb8 cost=$0.0074 (no caching: $0.0722) 4 calls: $0.1122 with caching vs $0.2886 without MISS key=da6706aa2f cost=$0.0902 (no caching: $0.0722)
Pause and think: From the output: across 4 calls we paid $0.1122 instead of $0.2886. Why is the first call more expensive than not caching at all?
The first call is a cache write, and in this pricing writes cost 1.25× normal input. We pay a small premium once so that later calls can read the prefix at 0.10×. Three hits more than repay it.
What we should put in the cache
- Long system prompts and behaviour rules that are the same for every user.
- Tool and function definitions for agents, which can run to thousands of tokens.
- Reference material: manuals, policies, codebases, long documents users ask many questions about.
- Few-shot examples that never change.
- Conversation history in multi-turn chats: each turn's prompt extends the previous one, so with a breakpoint near the end of the history every turn reuses the previous turn's cache.
Do not try to cache content that changes on every request, or short prompts below the minimum length. And remember caching is per exact prefix: two users with different system prompts get separate caches.
Benefits, real-world use and pitfalls
- Lower cost. Cached input tokens are billed at a large discount.
- Lower latency. Skipping prefill on a long prefix cuts time to first token, often substantially for very long prompts.
- Same output. The cached keys and values are the same numbers prefill would have produced, so caching is meant to leave answer quality unchanged.
- Enables long-context designs. Putting a whole manual or codebase in the prompt becomes affordable when it is paid for mostly once.
Prompt caching in the real world Coding agents resend a large system prompt, tool list and file contents on every step, so they lean heavily on caching. Chat products place a cache breakpoint at the end of the conversation so each new turn reuses the previous one. Document assistants cache a long contract once and answer many questions about it within minutes. Self-hosted servers like vLLM and SGLang reuse shared prefixes across users who share a system prompt.
Common mistakes Putting dynamic content (dates, user IDs, retrieved chunks) before the stable content. Serializing JSON or tools in a different order each time. Assuming a cache survives an hour when the TTL is 5 minutes and traffic is sparse. Not checking the usage fields in the API response, which report how many tokens were read from or written to the cache: if cached tokens show 0, the prefix is not matching.
Worked example, step by step
Chats are where caching feels most natural, so let us count one. A support chat has a 2,000-token system prompt. Each user message is 100 tokens and each reply is 300 tokens. We place the cache point at the end of every request. All numbers are illustrative.
| Turn | Input tokens | Read from cache | New |
|---|---|---|---|
| 1 | 2,000 + 100 = 2,100 | 0 | 2,100 |
| 2 | 2,100 + 300 + 100 = 2,500 | 2,100 | 400 |
| 3 | 2,500 + 300 + 100 = 2,900 | 2,500 | 400 |
Notice that the last reply counts as new input on the next turn. The model wrote it as output. When we send it back, it is fresh prompt text that sits after the cached point.
Pricing the three turns (write 1.25×, read 0.10×, in units of one normal input token)
- Without caching: 2,100 + 2,500 + 2,900 = 7,500 units.
- Turn 1: Everything is a write: 2,100 × 1.25 = 2,625.
- Turn 2: 2,100 read and 400 written: 210 + 500 = 710.
- Turn 3: 2,500 read and 400 written: 250 + 500 = 750.
- Total: 2,625 + 710 + 750 = 4,085 units, about 46% less than 7,500. The gap widens with every extra turn, because the read part keeps growing while the new part stays at 400.
The same table explains a classic failure. If the app edits an early message in the middle of a chat, for example by trimming the oldest turn or by re-wording the system prompt, the prefix changes at that point. Every token after it is new again, and that turn is billed like turn 1.
Practice: try it yourself
We will build a cache that works on the blocks of a prompt. For each request it finds the longest run of leading blocks it has seen before, and it forgets entries that were not used for 300 seconds. Token counts are illustrative.
practice_block_cache.py
import hashlib
TTL = 300 # seconds an unused entry survives
cache = {} # prefix hash -> time of last use
def key(blocks):
return hashlib.sha256("|".join(blocks).encode()).hexdigest()[:8]
def request(blocks, sizes, now):
# Find the longest prefix of blocks that is cached and not expired
hit = 0
for i in range(len(blocks), 0, -1):
k = key(blocks[:i])
if k in cache and now - cache[k] <= TTL:
hit = i
break
# Store (or refresh) every prefix of this request for later calls
for i in range(1, len(blocks) + 1):
cache[key(blocks[:i])] = now
read, fresh = sum(sizes[:hit]), sum(sizes[hit:])
print(f"t={now:3}s cached blocks={hit} read={read:5} computed={fresh:5}")
sizes = [4000, 20000, 50] # illustrative token counts per block
system, manual = "system rules v1", "product manual v7"
request([system, manual, "Q: how do I descale?"], sizes, now=0)
request([system, manual, "Q: is the lid dishwasher safe?"], sizes, now=60)
request([system, manual, "Q: what is the warranty?"], sizes, now=500)
request(["time 10:42", system, manual], [10, 4000, 20000], now=510)
request(["time 10:43", system, manual], [10, 4000, 20000], now=520)Output:
t= 0s cached blocks=0 read= 0 computed=24050 t= 60s cached blocks=2 read=24000 computed= 50 t=500s cached blocks=0 read= 0 computed=24050 t=510s cached blocks=0 read= 0 computed=24010 t=520s cached blocks=0 read= 0 computed=24010
Now change it:
- Change the third request from
now=500tonow=350. Predict hit or miss before running. Remember that each use resets the clock. - In the last two requests, move the time block to the end:
[system, manual, "time 10:42"]with sizes[4000, 20000, 10]. Predictreadandcomputedfor both. - Give the second request a new manual,
"product manual v8". Predictcached blocks,readandcomputedfor it.
Pause and think: Request 3 arrives at t=500 and misses, although the same prefix was a hit at t=60 and no text changed. What happened? And what would one question every 4 minutes have done?
The entry expired. It was last used at t=60, and 440 seconds passed, which is more than the 300-second TTL. Each use resets the clock, so a question every 240 seconds would have kept the entry alive without end. Sparse traffic caused this miss, not changed text. That is why low-traffic apps see fewer hits than their prompts suggest.
Pause and think: Our cache stores every leading run of blocks, not just the whole prompt. Suppose we stored one entry per request, keyed on all its blocks together. What would happen to request 2?
It would miss. Request 2 ends with a different question, so its full-prompt key was never stored. The saving comes from matching a prefix that is shorter than the whole prompt: system plus manual. This is why the stable part must end before the changing part begins, and why real APIs cache up to a marked point or in blocks instead of whole prompts.
Key takeaways
- Prefill computes keys and values for every prompt token; prompt caching stores them for a prefix and reuses them.
- Causal attention makes this safe: a prefix's KV never depends on what comes after it.
- Hits need an exact token match from the first token, so put stable content first and dynamic content last.
- Writes may cost a little extra, reads cost far less, and entries expire after a TTL refreshed by hits.
- Caching lowers cost and time to first token without changing the answer.
Key terms
- Prefill: The phase where the model processes all prompt tokens and builds the KV cache.
- KV cache: Stored key and value vectors for processed tokens, reused during generation.
- Prompt caching: Keeping the KV cache of a prompt prefix so later requests with the same prefix skip its prefill.
- Exact-prefix rule: A cache hit requires the prompt to match token-for-token from the start up to the cached point.
- Cache write / cache read: Storing a new prefix on a miss vs reusing a stored prefix on a hit; often priced differently.
- TTL: Time to live: how long an unused cache entry is kept before it expires.
- Time to first token (TTFT): Delay between sending a request and receiving the first generated token.
← 9.2 Prompt Chaining: Decomposing Complex Tasks into Steps · 9.4 Context Engineering: Curating the Model's Working Memory →