Modern AI Engineering

Lesson 10.9 · 22 min

Embedding Caches: Avoiding Redundant Embedding Calls

Every night our RAG pipeline re-embeds 2 million chunks, and 98% of them have not changed since yesterday. Why are we paying for the same numbers again?

In short: An embedding cache stores the vector computed for a piece of text so the same text never has to be embedded twice. The cache key is a hash of the exact text together with the model name (and version), the value is the vector, and a lookup either hits (return the stored vector) or misses (call the model, store the result). Eviction rules such as LRU and TTL keep it bounded, and it can live in memory, on disk or in a shared store like Redis.

What is an embedding, and how do we get one?

An embedding is a list of numbers (a vector), for example 768 or 1,536 floating-point values, that represents the meaning of a piece of text. Texts with similar meaning get vectors that are close together. Embeddings power semantic search, RAG, clustering, deduplication and recommendations.

To get one, we send text to an embedding model: either a hosted API (we pay per token and wait for a network round trip) or a model we run ourselves on a CPU or GPU (we pay in hardware and compute time). Either way, each call takes real time (from a few milliseconds locally to tens or hundreds of milliseconds over the network) and real money or compute.

One property makes caching possible: for a fixed model, embedding is (for practical purposes) deterministic. The same text through the same model version gives the same vector. Tiny floating-point differences can occur across hardware or batch sizes, but they are far too small to matter for search. So if we have computed it once, we can reuse it.

What is an embedding cache, and why do we need one?

An embedding cache is a key-value store that remembers "this text, with this model, gives this vector". Before calling the embedding model, we look in the cache. If the vector is there (a cache hit), we use it immediately. If not (a cache miss), we call the model, then save the result for next time.

Think of it like a translator's notebook A translator who keeps a notebook of sentences already translated never translates "Thank you for your order" twice. They look it up in a second. The notebook must note which language pair each entry is for, or a French translation might be handed out for a German request. In an embedding cache, the model name plays that role.

Where does repeated text come from? More places than we expect:

  • Re-indexing: nightly or on-deploy pipelines re-process whole document collections where most chunks are unchanged.
  • Popular queries: many users ask "how do I reset my password" with exactly the same words.
  • Duplicate content: boilerplate footers, legal disclaimers, repeated FAQ answers across pages.
  • Experiments: re-running an evaluation or a chunking experiment over the same texts.
  • Retries and crashes: a pipeline that fails at 80% should not pay again for the first 80%.

The benefits are direct: lower cost (fewer paid tokens), lower latency (a memory lookup is microseconds; a network call is milliseconds), fewer rate-limit problems with hosted APIs, and faster re-indexing.

The core idea and the cache key: hash of text plus model

The heart of any cache is the key: the label we look things up by. For embeddings, the key must change whenever the vector could change, and stay the same otherwise. Two things determine the vector: the exact text and the exact model (name and version, plus any settings that affect output, such as the output dimension or a "query" vs "document" mode).

Texts can be long, so instead of using the raw text as the key, we use a hash: a function such as SHA-256 that turns any input into a fixed-length fingerprint (64 hexadecimal characters for SHA-256). The same input always gives the same hash; any change, even one character, gives a completely different hash; and accidental collisions are astronomically unlikely.

The most dangerous bug: forgetting the model in the key If the key is only hash(text), then after upgrading the embedding model the cache keeps returning old-model vectors. New queries are embedded with the new model, old documents come from the cache with the old model, and similarity search silently becomes garbage, because vectors from different models are not comparable. Always include the model identity (and version) in the key, or use a separate cache namespace per model.

Should we normalise text before hashing (trim spaces, lowercase)? Trimming whitespace is usually safe. Lowercasing is not, unless the model itself is case-insensitive, because "Apple" and "apple" may get different vectors. A good rule: hash exactly the string we send to the model.

Pause and think: Two chunks differ only by a trailing space: "Refunds take 5 days" vs "Refunds take 5 days ". Will they share a cache entry?

Not if we hash the raw strings: one character of difference produces a completely different SHA-256 hash, so it is a miss. If we strip whitespace before both hashing and embedding, they become the same string and share an entry.

The request flow: a hit and a miss

The same flow in words

  1. Hash: Compute the key from the model identity and the text.
  2. Check: If the key is present and fresh, it is a hit: return the vector and mark the entry as recently used.
  3. Compute on miss: Otherwise call the embedding model. In batch pipelines, collect all misses and embed them in one batched call.
  4. Store: Save the new vector with the time it was created.
  5. Evict if needed: If the cache is over its size limit, remove the entry chosen by the eviction policy.

For indexing pipelines the batch version matters: hash all chunks, look up all keys at once, send only the misses to the model in batches, then write them back. If 98% of 2 million chunks are unchanged, we embed 40,000 chunks instead of 2 million.

Eviction: LRU and TTL

A cache cannot grow forever. A 1,536-dimension float32 vector is about 6 KB; 10 million of them is about 60 GB. Eviction decides what to throw away. Two rules are common, and often combined:

  • LRU (Least Recently Used): when the cache is full, remove the entry that has gone longest without being used. Popular texts stay; one-off texts fall out. Implemented with a hash map plus a usage-ordered list (Python's OrderedDict does both).
  • TTL (Time To Live): each entry expires after a fixed time, e.g. 1 hour or 30 days. After that it is treated as missing and recomputed. TTL bounds how long stale data can survive, for example if a model is silently updated behind the same name.

Other policies exist, such as LFU (Least Frequently Used, evict the entry with the fewest hits) and size-based limits. For embeddings of a stable model, entries never become "wrong" as long as the key includes the model version, so many teams use long TTLs or none at all, and rely on LRU or disk space limits.

embedding_cache.py

import hashlib
from collections import OrderedDict
MODEL = "embed-model-v2"         # placeholder model name
api_calls = 0
def slow_embed(text):            # stand-in for a paid embedding API call
global api_calls
api_calls += 1
return [len(text) / 100, text.count(" ") / 10]
class EmbeddingCache:
def __init__(self, capacity=3, ttl=3600):
self.data, self.capacity, self.ttl = OrderedDict(), capacity, ttl
def key(self, text):         # same text + same model -> same key
return hashlib.sha256(f"{MODEL}|{text}".encode()).hexdigest()[:12]
def get(self, text, now):
k = self.key(text)
if k in self.data:
if now - self.data[k][1] < self.ttl:
self.data.move_to_end(k)          # mark as recently used
return self.data[k][0], "hit", ""
status = "stale"                      # too old: recompute
else:
status = "miss"
vec = slow_embed(text)
self.data[k] = (vec, now)
self.data.move_to_end(k)
evicted = ""
if len(self.data) > self.capacity:        # full: drop least recently used
evicted = "evicts " + self.data.popitem(last=False)[0]
return vec, status, evicted
cache = EmbeddingCache()
events = [(0, "reset password"), (5, "refund policy"), (9, "reset password"),
(12, "track order"), (20, "store hours"), (30, "refund policy"),
(4000, "refund policy")]        # t=4000s: older than the 1h TTL
for t, text in events:
_, status, evicted = cache.get(text, now=t)
print(f"t={t:<5} {status:5} {text!r:17} key={cache.key(text)} {evicted}".rstrip())
print(f"requests={len(events)}  api_calls={api_calls}")

Output:

t=0     miss  'reset password'  key=b043a4f1675e
t=5     miss  'refund policy'   key=fac0cddb8a88
t=9     hit   'reset password'  key=b043a4f1675e
t=12    miss  'track order'     key=1b6e8ca0b362
t=20    miss  'store hours'     key=c03e18c5b18c evicts fac0cddb8a88
t=30    miss  'refund policy'   key=fac0cddb8a88 evicts b043a4f1675e
t=4000  stale 'refund policy'   key=fac0cddb8a88
requests=7  api_calls=6

Read the trace: at t=9 "reset password" is a hit. At t=20 the cache holds 3 entries, so adding "store hours" evicts the least recently used, "refund policy". At t=30 "refund policy" is a miss again and evicts "reset password". At t=4000 "refund policy" is in the cache but older than the 1-hour TTL, so it is stale and recomputed. 7 requests cost 6 calls here because the cache is tiny; real caches with thousands of entries and repetitive traffic reach much higher hit rates.

Pause and think: In the trace, why was "refund policy" evicted at t=20 even though "reset password" was added earlier?

LRU evicts the least recently used, not the oldest added. "reset password" was used again at t=9, so it was fresher than "refund policy", last used at t=5.

Where the cache lives: memory or disk

Storage size matters. Vectors can be stored as float32 (4 bytes per number), float16 (2 bytes), or compressed. Store them in a compact binary form rather than as JSON text, which can be several times bigger.

Benefits and the embedding cache in the real world

Real-world use LangChain offers a CacheBackedEmbeddings wrapper that stores vectors in a key-value store, namespaced by model, keyed by a hash of the text. Teams commonly keep embeddings in Redis or in a database table keyed by content hash. Data pipelines use content hashes to skip unchanged documents entirely. Some vector databases and RAG frameworks offer similar "skip if unchanged" ingestion features; check the current docs.

An embedding cache is different from a semantic cache (the next lesson). An embedding cache reuses a vector only for the exact same text. A semantic cache reuses a whole LLM answer for a similar question, which is riskier and saves far more per hit.

Common mistakes Leaving the model or version out of the key. Hashing a different string from the one actually embedded (e.g. hashing before adding a prefix like "query: "). Storing vectors as JSON text and running out of memory. No size limit, so the cache grows forever. Caching per-user text containing personal data without thinking about retention and deletion rules.

Worked example, step by step

Is the cache worth building? Let us put numbers on the nightly job from the start of this lesson. The price and the speed below are illustrative; plug in your own.

The nightly re-index, in numbers

  1. Size of the job: 2,000,000 chunks × 300 tokens each = 600 million tokens for one full run.
  2. Cost with no cache: At an illustrative $0.02 per million tokens: 600 × 0.02 = $12 per night, about $360 per month.
  3. Cost with a cache: 98% of chunks are unchanged, so we embed 40,000 chunks, or 12 million tokens: $0.24 per night.
  4. Time: If the model embeds 2,000 chunks per second, a full run takes 1,000 seconds (about 17 minutes). The 40,000 misses take 20 seconds, plus the time to hash and look up 2 million keys.
  5. Storage: 2,000,000 vectors × 1,536 numbers × 4 bytes ≈ 12.3 GB as float32. That is too much to hold inside every worker process, so the cache belongs in a shared store.

The storage number explains the "two-level" pattern from the comparison above. A small in-memory map answers the hottest keys at once, and a big shared store holds everything else.

One caution before building this for query traffic: measure the hit rate first, as hits / (hits + misses) on real logs. Re-indexing repeats almost everything, so the cache pays off at once. User queries may repeat far less. If only 5% of queries are exact repeats, the cache removes only 5% of the calls.

Practice: try it yourself

The earlier code handled one text at a time. Now we build the batch version used by indexing pipelines: hash every chunk, find the misses, embed only those in one batched call, and write them back. We replay four nightly runs, including one where the embedding model is upgraded.

practice_batch_cache.py

import hashlib
calls = 0
def embed_batch(texts):                 # stand-in for one batched model call
global calls
calls += len(texts)
return [[len(t), t.count(" ")] for t in texts]
def key(model, text):                   # model identity is part of the key
return hashlib.sha256(f"{model}|{text}".encode()).hexdigest()
def index(chunks, model, cache):
todo = {}                           # unique misses: key -> text
for c in chunks:
if key(model, c) not in cache:
todo[key(model, c)] = c
for k, vec in zip(todo, embed_batch(list(todo.values()))):   # embed misses only
cache[k] = vec
return [cache[key(model, c)] for c in chunks], len(todo)
monday = ["Refunds take 5 days.", "Shipping takes 2 days.", "Reset your password.",
"Contact support by chat.", "Refunds take 5 days."]      # note the duplicate
tuesday = monday[:3] + ["Contact support by chat or phone."]       # one chunk edited
cache = {}
runs = [("Mon, model v1", monday, "embed-v1"), ("Tue, model v1", tuesday, "embed-v1"),
("Wed, model v1", tuesday, "embed-v1"), ("Thu, model v2", tuesday, "embed-v2")]
for label, chunks, model in runs:
_, embedded = index(chunks, model, cache)
saved = 1 - embedded / len(chunks)
print(f"{label}: {len(chunks)} chunks, embedded {embedded}, calls saved {saved:.0%}")
print(f"total model calls = {calls}, cache entries = {len(cache)}")

Output:

Mon, model v1: 5 chunks, embedded 4, calls saved 20%
Tue, model v1: 4 chunks, embedded 1, calls saved 75%
Wed, model v1: 4 chunks, embedded 0, calls saved 100%
Thu, model v2: 4 chunks, embedded 4, calls saved 0%
total model calls = 9, cache entries = 9

Now change it:

  • Add a trailing space to one of the Tuesday chunks. Predict how many chunks are embedded on Tuesday now.
  • Change key so it ignores the model: hash only text. Predict the Thursday line, and explain why the "100% saved" it reports is bad news.
  • Add a fifth run, ("Fri, model v1", monday, "embed-v1"). Monday contains the old "Contact support by chat." chunk that was edited on Tuesday. Predict how many chunks are embedded on Friday.

Pause and think: On Thursday all 4 chunks were embedded again although no text changed. Is the cache broken?

No, it is doing its job. The model name is part of the key, so under model v2 none of the keys exist yet and every chunk is a miss. That is what keeps old-model vectors out of the new index. The v1 entries are still stored (9 entries in total) and now only waste space until they are evicted or their namespace is cleared.

Pause and think: On Monday the cache started empty, and the list had 5 chunks, but only 4 were embedded. The cache could not have helped yet. What saved the fifth call?

Removing duplicates inside the batch. Two chunks had the same text, so they had the same key, and the dictionary of misses kept one entry for both. Without that step, a cold cache would send both copies to the model, because the lookup for each happens before either result is stored.

Key takeaways

  • Embedding the same text with the same model gives the same vector, so we can compute it once and reuse it.
  • Key = hash(model + version + exact text); leaving the model out silently corrupts search after an upgrade.
  • A hit returns the stored vector; a miss calls the model and stores the result; batch the misses in pipelines.
  • LRU keeps the cache bounded by evicting the least recently used entry; TTL expires entries after a set time.
  • Caches can live in memory, on disk or in a shared store like Redis; they cut cost, latency and re-index time.

Key terms

  • Embedding cache: A key-value store that maps (model, text) to its previously computed vector.
  • Cache hit / miss: Whether the requested key is already in the cache (hit) or must be computed (miss).
  • Hash (SHA-256): A function that turns any input into a fixed-length fingerprint that changes completely if the input changes.
  • LRU: Least Recently Used eviction: remove the entry unused for the longest time.
  • TTL: Time To Live: an entry expires a fixed time after it was stored.
  • Cache namespace: A separate key space, e.g. one per embedding model, so entries never mix.

← 10.8 HyDE: Generating Hypothetical Documents to Improve RAG · 10.10 Semantic Caching: Skipping the LLM for Similar Queries →