Lesson 9.5 · 21 min
Context Compaction: Fitting More Into a Finite Window
How can an AI agent keep working for hours when its context window fills up in minutes?
In short: Every LLM has a context window: a fixed maximum number of tokens it can see at once. Long conversations and agent sessions eventually overflow it. Context compaction solves this by replacing older parts of the history with a compact summary that keeps the important facts, decisions and open tasks, while keeping the most recent turns word for word.
LLMs, context and the context window
A large language model (LLM) is a neural network that reads text as tokens (small pieces of words) and predicts the next token, over and over, to write a response. It has no memory of its own between calls. Each call, it sees only what we send.
That “what we send” is the context: the system prompt (standing instructions), the conversation so far, tool definitions, tool results and the new message. In a chat app, each new turn resends the whole history, because that is the only way the model can “remember” earlier turns.
The context window is the maximum number of tokens the model can process in one call. Modern models have windows from tens of thousands to around a million tokens, depending on the model. Input plus output must fit inside it. If the context is too long, the API rejects the request or the application must cut something.
Think of it like a whiteboard in a long meeting A team plans a project on one whiteboard. After two hours it is full. Erasing the oldest corner would wipe out the budget they agreed on. Instead, someone writes a neat box in the corner: “Decided: budget 4,000, dates 12–14 June, team of 8”, then erases the messy notes around it. The meeting continues with the key decisions still visible. That box is a compaction.
The problem of long conversations
Our running example: an assistant helping plan a team offsite. Over many turns, the user shares the team size, the budget, preferences, dates and a dietary need. Each turn adds tokens. An agent is even hungrier: every tool call adds its output, and a single web page or log file can be thousands of tokens.
- Hard limit. Eventually the history exceeds the window and the call fails.
- Rising cost and latency. Every turn resends everything before it, so cost per turn keeps growing even before the limit.
- Falling quality. Long, cluttered contexts make it harder for the model to focus. Key facts get buried in the middle among stale tool outputs.
The naive fix, and why it fails
The simplest fix is truncation: when the history is too long, drop the oldest messages. A common version is a sliding window that always keeps only the last N turns. It is easy and cheap, but it forgets blindly.
In our offsite chat, the very first assistant turn recorded “team of 8, budget 4,000”. A sliding window that keeps the last 4 turns drops it. A few turns later the assistant suggests a lodge for 20 people at 9,000 dollars, because it no longer knows the constraints. Old messages are often exactly where the goal, the rules and the key decisions live.
Pause and think: Pause and predict: why not just keep the first few turns and the last few turns, and drop the middle?
It helps a bit (the original goal survives), but facts stated in the middle, such as the dates or the vegetarian guest, are still lost. What we need is to keep the important information from every part of the history, not to keep certain positions. That requires understanding the content, which is what a summary does.
What is context compaction?
Context compaction is replacing a long stretch of context with a shorter version that keeps what matters for the rest of the task, then continuing from the compacted context. Usually the older part of the history is summarized and the most recent turns are kept verbatim, because they carry the immediate thread of the conversation.
Compaction has three parts: a trigger (when to compact, for example when the context passes 80% of the window), a method (how to shrink: summarize, trim tool outputs, drop duplicates) and a policy (what to keep raw: the system prompt, the last few turns, pinned facts).
How summarization powers compaction
The summarizer is usually the same LLM (or a cheaper one) called with a special prompt. The quality of the compaction is the quality of that prompt. A good compaction prompt asks the model to preserve, in a structured form:
- The goal the user is trying to achieve.
- Hard constraints and decisions: numbers, dates, names, choices already made (team of 8, budget 4,000, 12–14 June).
- Progress so far: what is done, what was tried and failed, and why.
- Open tasks and next steps.
- Important references: file paths, IDs, URLs, so the agent can look details up again instead of keeping them in the context.
It should drop chit-chat, pleasantries, repeated content and raw tool outputs whose conclusions are already captured. The summary is then inserted as a single message (often labelled as a summary of earlier conversation) in place of the old turns.
A step-by-step walkthrough
One compaction cycle
- Measure: After each turn, count the tokens in the context (most APIs return usage counts; a tokenizer works offline).
- Trigger: If the count crosses the threshold (say 80% of the window), start compaction before the next call fails.
- Split: Divide the history into an old part to summarize and a recent part (last K turns) to keep verbatim.
- Summarize: Send the old part to the LLM with a compaction prompt that asks for goals, facts, decisions and open tasks.
- Replace: Build the new context: system prompt + summary + recent turns. The old turns are removed (and can be archived on disk).
- Continue: The conversation goes on with plenty of room. The cycle repeats if the context fills up again; later summaries fold in the earlier summary.
Context compaction in code
A runnable toy version. The summarizer is a stand-in that keeps sentences tagged FACT; a real system would call an LLM with a compaction prompt. Tokens are counted as words.
compaction.py
def tokens(text):
return len(text.split()) # rough stand-in for a tokenizer
def summarize(turns):
# Stand-in for an LLM summary call: keep the facts, drop the chatter
facts = [t.split("FACT", 1)[1].strip() for t in turns if "FACT" in t]
return "SUMMARY: " + "; ".join(facts)
LIMIT, KEEP_RECENT = 50, 2 # compact when over 50 tokens; keep last 2 turns raw
turns = [
"user: hi, can you help me plan a small team offsite please",
"assistant: sure! FACT team of 8 people, budget 4000 dollars",
"user: we like hiking and good food, nothing too fancy",
"assistant: FACT likes hiking and food, casual style",
"user: FACT dates must be 12-14 June",
"assistant: great, I will look for lodges near trails for those dates",
"user: FACT one person is vegetarian",
"assistant: noted, I will check every menu for vegetarian options",
]
history = []
for t in turns:
history.append(t)
used = sum(tokens(h) for h in history)
if used > LIMIT: # trigger
old, recent = history[:-KEEP_RECENT], history[-KEEP_RECENT:]
history = [summarize(old)] + recent # replace old turns
after = sum(tokens(h) for h in history)
print(f"compacted {len(old)} turns: {used} -> {after} tokens")
print("\nContext the model sees next:")
for h in history:
print(" ", h)Output:
compacted 4 turns: 59 -> 33 tokens Context the model sees next: SUMMARY: team of 8 people, budget 4000 dollars; likes hiking and food, casual style user: FACT dates must be 12-14 June assistant: great, I will look for lodges near trails for those dates user: FACT one person is vegetarian assistant: noted, I will check every menu for vegetarian options
Pause and think: In the output, the team size and budget came from turn 2, which a 4-turn sliding window would have dropped. Where are they now?
Inside the SUMMARY line at the top of the context. Compaction kept the facts and dropped the chatter (“hi, can you help me…”, “nothing too fancy”), so the context shrank from 59 to 33 tokens without losing the constraints.
Compaction in real AI agents, and why it matters
Where you will meet it Coding agents such as Claude Code offer a manual compact command and also compact automatically when the conversation approaches the context limit, replacing the history with a summary of the work so far. Other agent tools and frameworks provide similar summarization or “memory” features, and some APIs offer server-side helpers that clear old tool results or summarize history. Details differ by product and change quickly, but the core idea is the same.
- Long tasks become possible. An agent can work through hundreds of tool calls instead of stopping when the window is full.
- Cost and latency stay bounded. The context size oscillates within a range instead of growing forever.
- Focus improves. Removing stale logs and dead ends leaves the model with a clean statement of goal and progress.
- Complementary tricks: clearing old tool outputs (keep only their conclusions), writing notes to a file the agent can re-read, and handing sub-tasks to sub-agents with fresh windows.
Compaction is lossy A summary can drop a detail that turns out to matter later (an exact error message, a specific number) or even state something slightly wrong, and the model will trust the summary. Mitigations: tell the summarizer to keep numbers, names, paths and decisions verbatim; keep recent turns raw; archive the full history so tools can look things up; avoid compacting in the middle of a delicate step; and test compaction on real long sessions.
Common mistakes and how to spot them
Compaction usually fails quietly. The session keeps running, and the damage shows up many turns later. These are the patterns to watch for.
| Mistake | What we see later | Fix |
|---|---|---|
| Summary of a summary | A fact survives the first compaction and is gone after the third | Tell the summarizer to carry the old summary forward unchanged and only add to it |
| Trigger set too late | The summarization call itself fails or is cut short | Trigger earlier, so the old turns plus the new summary still fit |
| Too few recent turns kept | The assistant repeats a step it has just done | Keep more recent turns word for word |
| Compacting in the middle of a step | A half-finished tool call is described vaguely, then redone wrongly | Compact only at natural boundaries, after a step is complete |
The trigger deserves a small calculation. Take a 20,000-token window, a trigger at 16,000 and a summary of about 2,000 tokens (illustrative).
How much room does compaction itself need?
- What the summary call reads: The old turns: at most about 16,000 tokens.
- What it writes: A summary of about 2,000 tokens. Together that is 18,000, which fits in 20,000.
- Move the trigger to 19,500: Now the same call needs about 21,500 tokens. It does not fit, and the compaction fails exactly when we need it.
- The rule: The gap between the trigger and the limit is not waste. It is the working room that compaction needs.
A simple safety net is a list of pinned facts: a few exact strings, such as the budget and the dates, that code looks for after every compaction. If one is missing we know at once, not twenty turns later.
Practice: try it yourself
We will run a longer session that compacts twice. The second summary must fold in the first one. After each compaction, code checks that two pinned facts are still present in the context.
practice_repeat_compaction.py
def tokens(text):
return len(text.split()) # rough stand-in for a tokenizer
LIMIT, KEEP_RECENT = 40, 2
PINNED = ["budget 4000", "12-14 June"] # facts that must never be lost
def summarize(old):
# Stand-in for an LLM summary: keep the earlier summary and KEY lines only
prior = [t[9:] for t in old if t.startswith("SUMMARY: ")]
kept = [t.split("KEY ", 1)[1] for t in old if "KEY " in t]
return "SUMMARY: " + "; ".join(prior + kept)
def compact(history):
old, recent = history[:-KEEP_RECENT], history[-KEEP_RECENT:]
return [summarize(old)] + recent
turns = [
"user: KEY budget 4000 for a team of 8",
"tool: lodge search returned 14 rows of prices and photos and reviews and maps",
"user: KEY dates 12-14 June",
"assistant: I found three lodges near trails that fit",
"tool: menu page with forty dishes listed one by one in great detail",
"user: KEY one vegetarian guest",
"assistant: Pine Lodge has vegetarian options every day",
"user: please book Pine Lodge for us",
]
history, count = [], 0
for t in turns:
history.append(t)
used = sum(tokens(h) for h in history)
if used > LIMIT: # trigger
history = compact(history)
count += 1
after = sum(tokens(h) for h in history)
lost = [f for f in PINNED if f not in " ".join(history)]
print(f"compaction {count}: {used} -> {after} tokens | pinned facts lost: {lost}")
print(history[0])Output:
compaction 1: 50 -> 33 tokens | pinned facts lost: [] compaction 2: 46 -> 24 tokens | pinned facts lost: [] SUMMARY: budget 4000 for a team of 8; dates 12-14 June
Now change it:
- In
summarize, changeprior + keptto justkept. Predict what the second compaction reports underpinned facts lost. - Set
KEEP_RECENT = 4. Predict how many compactions run and whether the context gets back under the limit. - Remove the word
KEYfrom the dates turn, so it readsuser: dates 12-14 June. Predict the output. What does it say about trusting the summarizer to notice what matters?
Pause and think: The second compaction summarized three items: the first summary and two later turns. Neither of those turns had a KEY line. Why did the budget and the dates still survive?
Because summarize carries the earlier summary forward: a line that starts with SUMMARY: is copied into the new summary before any new key lines are added. Each compaction folds the previous one in. A summarizer that looked only for new key lines would have produced an empty summary here and wiped out everything learned in the first half of the session.
Pause and think: The pinned-fact check looks for exact strings such as budget 4000. A real LLM summarizer might write “a budget of 4,000 dollars”. Is the fact lost? What should we do?
The fact is still there, but the exact-string check would report it as lost. Two fixes work together. Tell the summarizer to copy numbers, dates and names word for word, which also protects them from being distorted. And keep the check strict, so that any re-wording of a critical value is flagged for a human to look at.
Key takeaways
- The context window is a hard token limit; long chats and agent sessions eventually overflow it.
- Truncation and sliding windows forget by position and can drop early goals and constraints.
- Compaction summarizes older history into goals, facts, decisions and next steps, and keeps recent turns raw.
- It needs a trigger, a summarization method and a policy for what to keep verbatim.
- Compaction is lossy: preserve exact values, archive the full history, and test it on real sessions.
Key terms
- Context: Everything sent to the model in one call: instructions, history, tool data and the new message.
- Context window: The maximum number of tokens a model can process in one call, input plus output.
- Truncation: Cutting old messages to make the context fit, without regard to their importance.
- Sliding window: Keeping only the last N messages of a conversation.
- Context compaction: Replacing older context with a compact summary so the session can continue within the window.
- Compaction trigger: The condition, such as a token threshold, that starts a compaction.
← 9.4 Context Engineering: Curating the Model's Working Memory · 10.1 Vector Databases: Storing and Searching Embeddings at Scale →