Modern AI Engineering

Lesson 10.10 · 23 min

Semantic Caching: Skipping the LLM for Similar Queries

"How do I get a refund?" and "Can I get my money back?" are different strings but the same question. Why pay an LLM twice to answer it?

In short: A semantic cache stores past LLM answers together with the embedding of the question that produced them. When a new question arrives, we embed it and look for a stored question whose embedding is similar enough (above a threshold); if found, we return the stored answer instead of calling the LLM. It saves cost and latency on repeated questions, but the threshold must be tuned carefully, because a too-low threshold returns answers to questions that only look similar.

What is a cache, and why traditional caching fails for AI apps

A cache is a fast store of results we already computed, so we can reuse them instead of computing again. Web browsers cache images; databases cache query results. A cache maps a key to a value. On a hit (key found) we return the value instantly; on a miss we compute it and store it.

A traditional cache uses exact matching: the key is the exact request, often hashed. That works when the same request repeats byte for byte. But people ask LLM apps in their own words. Our running example, a shop's support chatbot, receives:

  • "How do I get a refund?"
  • "Can I get my money back?"
  • "how do i get a refund" (lowercase, no question mark)
  • "What's the process for refunds?"

An exact-match cache treats all four as different keys, so it calls the LLM four times and gets four nearly identical answers. LLM calls are the slowest and most expensive part of the app, often seconds and a noticeable cost per call, so this waste adds up quickly at scale.

Think of it like an experienced shop assistant A new assistant looks up every question in the manual. An experienced one recognises "can I get my money back?" as the refund question they answered ten times today and replies straight away, without caring about the exact words. A semantic cache gives an app that kind of memory.

What is semantic caching? Embeddings and similarity

Semantic caching matches questions by meaning instead of exact text. "Semantic" means "about meaning". We store each answered question as a vector, and when a new question's vector is close enough to a stored one, we reuse that stored answer.

To compare meanings, we use embeddings: lists of numbers (vectors) produced by an embedding model, arranged so that texts with similar meaning get vectors pointing in similar directions. Two paraphrases of the refund question get vectors that are very close; a question about shipping gets a vector pointing elsewhere.

Similarity between embeddings is usually measured with cosine similarity: the cosine of the angle between two vectors. It is 1 when they point the same way, near 0 when unrelated. For vectors of length 1, it equals the dot product.

How semantic caching works step by step

Handling one request

  1. Embed the question: Turn the incoming question into a vector with an embedding model (a fast, cheap call compared with the LLM).
  2. Search the cache: Find the most similar stored question vector, usually with a vector index for speed.
  3. Compare with the threshold: If the best similarity is at or above the threshold (e.g. 0.95), it is a hit.
  4. Hit: return the stored answer: Skip the LLM entirely. Latency drops from seconds to milliseconds.
  5. Miss: call the LLM and store: Generate a fresh answer, return it, and save (question vector, question, answer) for the future.

The cache usually lives in a vector database or a vector-capable store such as Redis, alongside metadata: when the entry was created, which model produced the answer, which user or tenant it belongs to, and a time-to-live (TTL) after which it expires.

A numeric walkthrough in code

This script runs six support questions through a semantic cache with threshold 0.95. The embeddings are hand-made 4-number vectors whose axes mean roughly refund, shipping, password and cancel, so we can read them; a real system would get 768+ numbers from an embedding model.

semantic_cache.py

import numpy as np
def unit(v):
v = np.array(v, dtype=float)
return v / np.linalg.norm(v)
# Pretend embeddings: axes = [refund, shipping, password, cancel]
E = {"How do I get a refund?":          unit([.9, .1, 0, .2]),
"Can I get my money back?":        unit([.85, .15, .05, .25]),
"How long does shipping take?":    unit([.1, .95, 0, .05]),
"When will my order arrive?":      unit([.1, .9, 0, .15]),
"How do I cancel my order?":       unit([.35, .3, 0, .9]),
"How do I cancel my subscription?": unit([.4, .05, .1, .9])}
cache = []                                 # list of (embedding, question, answer)
THRESHOLD = 0.95
def ask(q):
v = E[q]
if cache:
sims = [float(v @ e) for e, _, _ in cache]
best = int(np.argmax(sims))
if sims[best] >= THRESHOLD:
return f"HIT  {sims[best]:.3f} -> reuse answer to {cache[best][1]!r}"
answer = f"<LLM answer for {q!r}>"     # the expensive call
cache.append((v, q, answer))
return "MISS -> call LLM, store answer"
for q in E:
print(f"{q:34} {ask(q)}")

Output:

How do I get a refund?             MISS -> call LLM, store answer
Can I get my money back?           HIT  0.994 -> reuse answer to 'How do I get a refund?'
How long does shipping take?       MISS -> call LLM, store answer
When will my order arrive?         HIT  0.994 -> reuse answer to 'How long does shipping take?'
How do I cancel my order?          MISS -> call LLM, store answer
How do I cancel my subscription?   HIT  0.963 -> reuse answer to 'How do I cancel my order?'

Read the results. "Can I get my money back?" is a correct hit (0.994) on the refund question. "When will my order arrive?" is a correct hit on the shipping question. But look at the last line: "How do I cancel my subscription?" hits "How do I cancel my order?" with similarity 0.963. These are different questions with different answers. This is a false hit: the cache confidently returns the wrong answer. With a threshold of 0.97 it would have been a miss.

Setting the similarity threshold

The threshold is the minimum similarity that counts as "the same question". It is the most important setting in a semantic cache, and it is a trade-off:

  • Too low (e.g. 0.80): many hits, big savings, but more false hits, where users get answers to a different question. This is the dangerous failure.
  • Too high (e.g. 0.99): almost no false hits, but true paraphrases are missed, so the cache barely saves anything.

How to choose: collect real question pairs from logs, label each pair "same answer" or "different answer", compute their similarities with your actual embedding model, and pick the threshold that keeps false hits below what you can tolerate. Similarity values are model-specific: 0.90 from one model is not the same as 0.90 from another. Teams often start conservatively (high) and lower it gradually while monitoring.

Pause and think: In our code, which threshold would have avoided the false hit while keeping both correct hits?

Anything above 0.963 and at or below 0.994, for example 0.97 or 0.98. The cancel-order/cancel-subscription pair scored 0.963, while the true paraphrase pairs scored 0.994.

Advantages of semantic caching

The benefits: lower cost (each hit avoids an LLM call), lower latency (milliseconds instead of seconds), less load on rate-limited APIs, and more consistent answers to the same question. Savings depend entirely on how repetitive the traffic is: support bots and FAQ assistants repeat a lot; open-ended creative or highly personal chats repeat little.

Real-world tools Open-source projects such as GPTCache (from Zilliz) and the semantic cache in RedisVL implement this pattern, and several LLM gateways and frameworks (for example LangChain's cache integrations) offer semantic caching options. Do not confuse it with prompt caching offered by LLM providers, which reuses the model's internal computation for an identical prompt prefix and still generates a fresh answer.

Things to keep in mind

Where semantic caches go wrong False hits on questions that differ in a small but crucial detail: "cancel my order" vs "cancel my subscription", "flights to Paris" vs "flights from Paris", "Python 2" vs "Python 3". Personal or account-specific answers: "what is my balance?" must never be served from another user's cache entry; scope entries by user or tenant, or do not cache them. Stale answers: prices and policies change; use a TTL and clear entries when source documents change. Conversation context: "and what about the blue one?" depends on earlier turns, so caching it on its own text is wrong.

  • Cache what is stable and shared: FAQs, policy questions, how-to questions.
  • Do not cache time-sensitive, personal, or context-dependent questions, or answers that relied on tool calls with live data.
  • Store metadata: model version, prompt version, creation time; invalidate when any of them change.
  • Monitor: log hits with their similarity, sample them for review, and track user feedback on cached answers.
  • Mind the overhead: every request now pays for an embedding and a vector search; with very low hit rates the cache can cost more than it saves.

Pause and think: Our bank chatbot caches "What is my current balance?" and a second user asking the same thing receives the first user's balance. What went wrong, and how do we fix it?

The answer is personal, so it must not be shared between users. Either exclude personal or account-specific questions from the cache, or scope cache entries by user (include the user ID in the lookup), and avoid caching answers that came from live account data.

Semantic caching is a simple, powerful idea: remember answers by meaning. It works best when many people ask the same stable questions in different words, with a carefully tuned threshold and clear rules about what may be cached.

Worked example, step by step

A semantic cache is not free. Every request pays for an embedding and a vector search, whether it ends in a hit or not. So when does the cache pay for itself? Let us work it out with illustrative prices and timings.

Finding the break-even point

  1. Write down the costs: One LLM call: $0.01 and 2,000 ms. One cache check (embedding plus search): $0.0001 and 30 ms.
  2. Cost per request: With a cache: check + (1 − h) × LLM, where h is the hit rate. Without a cache: just the LLM cost.
  3. Break-even: The cache wins when h × LLM cost is more than the check cost: h > 0.0001 / 0.01 = 1%.
  4. A support bot, h = 30%: Per 10,000 requests: no cache costs $100. With the cache: $1 for checks + 7,000 × $0.01 = $71. We save $29.
  5. A creative-writing app, h = 0.5%: $1 for checks + 9,950 × $0.01 = $100.50. The cache now costs more than having none.
  6. Latency: Average = 30 + (1 − h) × 2,000 ms. At h = 30% that is 1,430 ms instead of 2,000. Note that every miss got 30 ms slower.
Per 10,000 requests, with the illustrative prices above.
Hit rateTotal costAverage latency
No cache$100.002,000 ms
0.5%$100.502,020 ms
10%$91.001,830 ms
30%$71.001,430 ms
60%$41.00830 ms

One cost is missing from the table: false hits. Suppose a wrong answer leads to a support escalation worth $5 (illustrative). At h = 30% we serve 3,000 answers from the cache. If 1% of them are false hits, that is 30 wrong answers, or $150 of damage, against $29 saved. This is why we fix the threshold for safety first and only then look at the hit rate.

Practice: try it yourself

The lesson said: label real question pairs, then pick the threshold that keeps false hits low enough. Now we code that. We take 15 labelled pairs, sweep six thresholds, count good hits and false hits at each, and choose the lowest threshold that stays within our tolerance.

practice_threshold_sweep.py

# Labelled pairs from (pretend) logs: the similarity of a new question to its
# nearest cached question, and whether the cached answer really was the right one.
# Illustrative numbers.
pairs = [
(0.99, True), (0.98, True), (0.97, True), (0.96, False), (0.96, True),
(0.95, True), (0.94, False), (0.93, True), (0.91, False), (0.90, True),
(0.88, False), (0.86, False), (0.82, False), (0.75, False), (0.60, False),
]
MAX_FALSE_HITS = 0                  # how many wrong answers we tolerate in this sample
def evaluate(threshold):
hits = [same for sim, same in pairs if sim >= threshold]
good = sum(hits)                # True counts as 1
return len(hits) / len(pairs), good, len(hits) - good
best = None
print("threshold  hit rate  good hits  false hits")
for t in [0.85, 0.90, 0.93, 0.95, 0.97, 0.99]:
hit_rate, good, false_hits = evaluate(t)
print(f"   {t:.2f}     {hit_rate:6.0%}  {good:9}  {false_hits:10}")
if best is None and false_hits <= MAX_FALSE_HITS:
best = t                    # lowest threshold that is still safe
print("chosen threshold:", best)

Output:

threshold  hit rate  good hits  false hits
0.85        80%          7           5
0.90        67%          7           3
0.93        53%          6           2
0.95        40%          5           1
0.97        20%          3           0
0.99         7%          1           0
chosen threshold: 0.97

Now change it:

  • Set MAX_FALSE_HITS = 1. Read the table in the output and predict the chosen threshold before you run it.
  • Add one more pair to the list: (0.98, False). Predict the new chosen threshold and the hit rate we are left with.
  • Add 0.96 to the list of thresholds to try. Predict the number of false hits in that row.

Pause and think: At 0.95 there is 1 false hit and the hit rate is 40%. At 0.97 there are none and the hit rate is 20%. A teammate wants 0.95 "because it doubles the savings". What two things should we ask before agreeing?

First: what does one wrong answer cost compared with one LLM call? If a false hit can mislead a customer about refunds, a few saved cents do not cover it. Second: is the sample big enough? Fifteen pairs is tiny. Zero false hits in 15 does not prove the rate is zero, so we should label many more pairs before trusting either threshold.

Pause and think: The list contains (0.96, True) and (0.96, False): the same similarity with opposite labels. What does that tell us about what a threshold can and cannot do?

No threshold can separate those two pairs, because similarity is the only thing it looks at. To tell them apart we need another signal: scoping entries by user or product, checking that key terms match (order against subscription), or a second, more careful check on borderline hits.

Key takeaways

  • Exact-match caches miss paraphrases; semantic caches match questions by embedding similarity.
  • Flow: embed the question, find the nearest cached question, hit if similarity ≥ threshold, else call the LLM and store.
  • The threshold trades hit rate against false hits; tune it on labelled pairs from real logs, per embedding model.
  • Hits save LLM cost and cut latency from seconds to milliseconds, most of all for repetitive FAQ-style traffic.
  • Never cache personal, time-sensitive or context-dependent answers without scoping, TTLs and invalidation.

Key terms

  • Semantic cache: A cache that reuses stored LLM answers for new questions with similar meaning.
  • Similarity threshold: The minimum similarity between question embeddings that counts as a cache hit.
  • False hit: A cache hit that returns the answer to a different question.
  • Cosine similarity: The cosine of the angle between two vectors; close to 1 means very similar meaning.
  • TTL: Time to live: how long a cache entry stays valid before it expires.
  • Prompt caching: A provider feature that reuses computation for identical prompt prefixes; it still generates a fresh answer.

← 10.9 Embedding Caches: Avoiding Redundant Embedding Calls · 10.11 Agentic RAG: Dynamic Retrieval with Multi-Step Reasoning →