Modern AI Engineering

Lesson 11.7 · 25 min

Agent Memory: Short-Term, Long-Term, and Episodic

An LLM forgets everything the moment a call ends; so how does an assistant remember that you are vegetarian three weeks later?

In short: LLMs are stateless, so agent memory is something we build around the model. It forms a stack: the context window as working memory, the session history as short-term memory, and external stores as long-term memory (facts, past episodes and procedures). Memory systems run four core operations, write, read, update and forget, and the hard part is choosing what is worth remembering and retrieving only what helps the current task.

The big picture

Every LLM call starts from a blank slate. The model's weights hold general knowledge from training, but nothing about you, this task, or what happened five minutes ago unless we put it into the prompt. Each call is independent: this is what we mean when we say the model is stateless.

Yet agents seem to remember: they keep track of a long task, recall your preferences, and avoid repeating last week's mistakes. All of that is memory engineering: ordinary software that decides what to save, where to keep it, and what to put back into the prompt at the right moment.

Think of it like a doctor with patient files A doctor sees hundreds of patients and cannot remember them all. Before your appointment, they pull your file: allergies, past treatments, notes from last visit. During the visit they keep a few things in their head (working memory) and jot notes. Afterwards they update the file, correcting what changed and leaving out small talk. The doctor's skill is like the LLM's weights; the file system is agent memory.

Why AI agents need memory

Without memory, an agent hits four walls:

  • Within a task: a multi-step agent must remember earlier tool results. The agent loop does this by keeping the message history, but long tasks overflow the context window.
  • Across sessions: users expect not to repeat themselves (“I told you last week I'm vegetarian”).
  • Personalisation: preferences, roles, projects and past decisions make answers more useful.
  • Learning from experience: an agent that remembers “this API needs a date in ISO format” fails less often next time.

Why not just keep everything in the prompt? Because the context window (the maximum tokens a model can read at once) is finite, every extra token costs money and latency on every call, and models use information in very long contexts less reliably (for example, details buried in the middle can be overlooked). Memory is therefore about selection: the right few facts, not all facts.

Pause and think: A model has a 200,000-token context window. Does that mean we no longer need a memory system?

No. Context resets with every new session, so cross-session memory still needs storage. Even within a session, filling the window is slow and expensive on every call, and models may overlook details in very long contexts. Large windows reduce the pressure but do not replace selecting what matters.

The memory stack

It helps to picture memory as a stack of layers, from fastest and smallest to slowest and largest. The names below are widely used, though different frameworks slice them a little differently.

The agent memory stack
LayerWhat it holdsLifetimeWhere it lives
Parametric memoryGeneral knowledge learned in training.Fixed until retraining or fine-tuningModel weights
Working memoryEverything in the current prompt: instructions, recent messages, retrieved facts.One model callContext window
Short-term (session) memoryThe running conversation and tool results of this task.One session or taskMessage list, often trimmed or summarised
Long-term: semanticFacts and preferences: “User is vegetarian”, “Team uses PostgreSQL”.Until changed or deletedDatabase, key-value or vector store
Long-term: episodicRecords of past events: “On 3 May we debugged the payment timeout; fix was X”.Until expired or archivedLogs or vector store with timestamps
Long-term: proceduralHow to do things: rules, learned instructions, reusable skills.Until revisedSystem prompt, instruction files, skill libraries

The terms semantic, episodic and procedural come from human memory research and map nicely onto agents: what I know, what happened, and how to do it. Only working memory is directly visible to the model. Every other layer matters only when its content is copied into the context.

Short-term memory management When a session grows too long, common tricks are: keep the last N messages verbatim, replace older ones with a running summary, and drop or shorten bulky tool outputs while keeping their key results. The goal and any user constraints should always stay in view.

The four core operations

Whatever storage we use, a memory system does four things:

Reading usually combines several signals into a score. A simple and common recipe is a weighted sum of relevance (how similar the memory is to the current query, often by embedding cosine similarity), recency (newer memories count more) and sometimes importance (how significant the memory was judged when written).

How memory flows at runtime

The same flow, step by step

  1. Retrieve before thinking: Before the model call, query long-term memory with the user's message (and the current goal). Keep only the top few results above a relevance threshold.
  2. Assemble the context: Place memories in a clearly labelled section (for example “Known facts about the user”) so the model treats them as background, not instructions.
  3. Run the agent: The loop proceeds as usual. Tool results accumulate as short-term memory.
  4. Extract: After the turn (often in the background), an LLM call or rules pick out durable facts, decisions and lessons from the conversation.
  5. Reconcile: Compare each candidate with existing memories: add if new, update if it changes an old fact, skip if duplicate, delete if the user asked to forget.

Writing can happen in the hot path (the agent itself calls a save_memory tool during the conversation) or in the background (a separate process reviews conversations later). Hot-path writes are immediate and transparent; background writes keep responses fast and allow more careful consolidation.

Here is a runnable toy memory store with all four operations. It uses a bag-of-words vector as a stand-in for a real embedding model, and the scoring formula above.

agent_memory.py

import numpy as np
VOCAB = ["diet", "vegetarian", "meat", "city", "berlin", "munich", "dog", "allergy", "nuts", "food"]
def embed(text):                         # toy bag-of-words "embedding"
v = np.array([float(w in text.lower()) for w in VOCAB])
return v / (np.linalg.norm(v) or 1.0)
store = {}                               # long-term memory: key -> record
def write(key, text, day):  store[key] = {"text": text, "day": day, "vec": embed(text)}
def forget(key):            store.pop(key, None)
def read(query, today, k=2, half_life=30):
q = embed(query)
scored = []
for key, m in store.items():
sim = float(q @ m["vec"])                       # relevance
recency = 0.5 ** ((today - m["day"]) / half_life)   # older = weaker
scored.append((round(0.8 * sim + 0.2 * recency, 3), key))
return sorted(scored, reverse=True)[:k]
write("diet", "User diet: vegetarian, no meat", day=1)
write("home", "User city: Berlin", day=2)
write("allergy", "User allergy: nuts (food)", day=3)
write("pet", "User has a dog named Rex", day=4)
print("day 10 read 'food ideas, diet?':", read("food ideas, diet?", today=10))
write("home", "User city: Munich (moved)", day=40)  # UPDATE: same key overwrites
forget("pet")                                         # FORGET: user asked us to
print("day 41 read 'which city':", read("which city?", today=41, k=1))
print("stored keys:", sorted(store), "| home =", store["home"]["text"])

Output:

day 10 read 'food ideas, diet?': [(0.497, 'allergy'), (0.489, 'diet')]
day 41 read 'which city': [(0.761, 'home')]
stored keys: ['allergy', 'diet', 'home'] | home = User city: Munich (moved)

Pause and think: In the matrix, the pet memory has the highest recency (0.871) but a low score. Why?

Its similarity to the query is 0: no shared words with “food ideas, diet?”. With 80% of the weight on relevance, recency alone can only add up to 0.2. That is intended: recent but irrelevant memories should not crowd out relevant ones.

What to store and what not to store

A practical guide
StoreUsually do not store
Stable preferences: diet, language, units, tone.Small talk and pleasantries.
Facts about ongoing work: project names, tech stack, deadlines.Raw tool outputs and full documents (store a pointer or a summary instead).
Decisions and their reasons: “chose vendor B because of price”.Guesses the agent made that were never confirmed.
Lessons learned: “the billing API needs ISO dates”.Secrets: passwords, API keys, card numbers.
Explicit “please remember” requests.Sensitive personal data without consent and a clear purpose.

Privacy is a design requirement Memory turns a stateless model into a system that keeps personal data. Tell users what is remembered, let them view and delete it, scope memories per user (never leak one user's memory into another's session), follow data-protection rules in your region, and honour “forget this” requests with real deletion.

Real-world use Several consumer chat assistants now offer saved memories that users can view and delete. Coding agents commonly read project instruction files (procedural memory) at the start of each session. Customer-support agents pull the customer's account history and past tickets (episodic memory) before replying.

Common mistakes and how to fix them

Memory failures and fixes
MistakeSymptomFix
Storing everythingRetrieval returns noise; answers get worse over time.Write only durable, useful facts; extract and summarise before storing.
Never updatingContradictory memories (“lives in Berlin” and “lives in Munich”).Reconcile on write: detect conflicts by key or similarity and overwrite.
Never forgettingStale facts resurface months later.Expiry, recency weighting, user-facing delete.
Retrieving too muchPrompt bloats; key facts get lost in the middle.Top-k with a relevance threshold; keep memory sections short.
Treating memories as instructionsA stored note like “always approve refunds” changes behaviour.Label memories as data; validate what gets written; guard against memory poisoning.
No user scopingOne user's details appear in another's chat.Partition every store by user or tenant and enforce it in code.

When not to add long-term memory. For one-off tasks, anonymous users, or highly regulated data where retention is risky, a stateless agent with only session memory may be the right design. Memory adds complexity and responsibility; add it when personalisation or continuity clearly pays off.

Quick summary. LLMs are stateless; agent memory is built around them as a stack: weights, the context window, session history, and long-term stores of semantic, episodic and procedural memory. Four operations keep it healthy: write, read, update and forget. Retrieve a few relevant memories by similarity and recency, store only durable and safe information, and give users control.

Worked example, step by step

The flow section ended with a step called reconcile: compare each new candidate memory with what is already stored. It is the step most often skipped, so let us do it by hand. Our store holds three memories about one user.

  • diet: vegetarian (written on day 1)
  • city: Berlin (written on day 2)
  • allergy: nuts (written on day 3)

On day 40 the user chats with the agent. After the turn, an extractor pulls out five candidate facts. For each one we ask two questions: does a memory about the same thing already exist, and if so, does the new fact agree with it?

Reconciling five candidates against the store
Candidate from the chatExisting memoryDecisionWhy
“I moved to Munich”city: BerlinUpdateSame topic, different value. The new fact replaces the old one.
“I still don't eat meat”diet: vegetarianSkipSame topic, same meaning. Storing it again would create a duplicate.
“My dog is called Rex”NoneAddNew topic, stable, and plausibly useful later.
“Forget my allergy”allergy: nutsDeleteAn explicit request. The memory is removed, not just hidden.
“Maybe I'll try sushi”diet: vegetarianSkipA passing thought, not a confirmed change. Unconfirmed guesses are not stored.

What makes each decision possible

  1. A way to find the matching memory: With keys like city, matching is exact. With free-text memories, we search for the most similar stored memory and treat a close match as “same topic”.
  2. A way to compare meaning: “Vegetarian” and “doesn't eat meat” are different strings with the same meaning. Simple code cannot see that, so this comparison is usually given to an LLM.
  3. A rule for uncertainty: The sushi line could be read as a diet change. When a candidate is unsure, the safe choice is to skip it, or to ask the user.
  4. A timestamp on every write: The updated city memory gets day 40. Later, recency scoring and audits both depend on knowing when a fact was last confirmed.

After reconciling, the store holds diet: vegetarian, city: Munich and pet: dog named Rex. Five candidates came in; the store still has three entries, with no contradiction and no duplicate. That is the sign of a healthy memory.

Practice: try it yourself

The earlier code showed how to read memories with a score. Here we build the write side: a background job that looks at each user message, extracts candidate facts, and reconciles them with the store. The extractor is a scripted stand-in for an LLM. The reconcile logic is real.

practice_memory_reconcile.py

# Background memory writing: extract candidate facts, then reconcile with the store.
SCRIPT = {   # what an extractor LLM would pull out of each user message (scripted)
"I'm vegetarian and I live in Berlin.": [("diet", "vegetarian"), ("city", "Berlin")],
"Nice weather today!":                  [],
"As I said, I don't eat meat.":         [("diet", "vegetarian")],
"I moved to Munich last week.":         [("city", "Munich")],
"Please forget where I live.":          [("city", None)],
"My password is tulip42, remember it.": [("password", "tulip42")],
}
def fake_extractor(message):
"""Stands in for the LLM call that finds durable facts in a message."""
return SCRIPT[message]
BLOCKED = {"password", "card_number"}      # things we never store
store = {}                                 # long-term memory: key -> record
def reconcile(key, value, day):
"""Compare one candidate with the store and pick exactly one decision."""
if key in BLOCKED:
return "REFUSE"
if value is None:                      # the user asked us to forget
return "DELETE" if store.pop(key, None) else "SKIP"
if key not in store:
store[key] = {"value": value, "day": day}
return "ADD"
if store[key]["value"] == value:       # we already know this
return "SKIP"
store[key] = {"value": value, "day": day}
return "UPDATE"
for day, message in enumerate(SCRIPT, start=1):
facts = fake_extractor(message)
decisions = [f"{reconcile(k, v, day)} {k}" for k, v in facts] or ["nothing durable"]
print(f"day {day}: {', '.join(decisions)}")
print("store:", store)

Output:

day 1: ADD diet, ADD city
day 2: nothing durable
day 3: SKIP diet
day 4: UPDATE city
day 5: DELETE city
day 6: REFUSE password
store: {'diet': {'value': 'vegetarian', 'day': 1}}

Now change it:

  • Change the day 3 fact to ("diet", "no meat"). Predict the decision for day 3. Is the result what a human would want? What would have to change in reconcile to get it right?
  • Remove "password" from BLOCKED. Predict the final store, then say which row of the “what not to store” table this breaks.
  • Swap the order of the day 4 and day 5 messages in SCRIPT (forget first, then the move). Predict both decisions and the final store. Is storing Munich after a forget request the right behaviour?

Pause and think: On day 3 the user repeats something we already know, and the decision is SKIP. A simpler design would just write it again. What goes wrong over months with “always write”?

The store fills with near-copies of the same fact. Retrieval then returns several duplicates in its top few results, pushing out other useful memories and wasting prompt space. Worse, when the fact later changes, an update may fix only one copy, leaving old copies to contradict it. Skipping duplicates at write time is far cheaper than cleaning them up later.

Pause and think: In this toy, exact keys make matching easy. Real extractors often produce free text such as “user relocated to Munich”. Which earlier idea from this lesson would we use to find the memory it conflicts with, and what new risk does that bring?

We would embed the candidate and search the store by similarity, as in the read operation, then treat a close match as the same topic. The risk is a wrong match. If “user's sister lives in Munich” is judged similar to “user lives in Berlin”, an update would overwrite a true fact with a wrong one. So similarity finds the candidates, and a careful comparison (often an LLM call) should make the final decision.

Key takeaways

  • LLMs are stateless; memory is software that decides what to put back into the context.
  • The stack: weights, working memory (context), session memory, and long-term semantic, episodic and procedural memory.
  • Four operations keep memory useful: write, read, update and forget.
  • Retrieve a few relevant memories using signals like similarity and recency, not everything.
  • Store durable, useful, safe facts; scope per user, let users delete, and never treat memories as instructions.

Key terms

  • Stateless: Keeping no information between calls; each LLM call sees only what is in its prompt.
  • Working memory: The content of the context window for the current model call.
  • Semantic memory: Long-term facts and preferences, such as “the user is vegetarian”.
  • Episodic memory: Long-term records of specific past events or interactions, usually timestamped.
  • Procedural memory: Knowledge of how to do things, stored as instructions, rules or reusable skills.
  • Memory consolidation: Merging, correcting and summarising stored memories so they stay consistent and compact.

← 11.6 Reflection Agents: Self-Critique for Higher-Quality Outputs · 11.8 Model Context Protocol: A Standard Interface for Agent Tools →