Lesson 9.4 · 23 min
Context Engineering: Curating the Model's Working Memory
When an AI assistant gives a wrong answer, is the model usually the problem, or is it what we showed the model?
In short: Context engineering is the work of deciding what information goes into an LLM's context window for each call: instructions, memory, history, retrieved documents, tool definitions and tool results. The model only knows what is in front of it, so selecting, ordering, compressing and isolating that information often matters more than clever wording. Good context is relevant, sufficient, compact and well structured.
What is context engineering?
A large language model (LLM) has no memory between calls and cannot look anything up by itself. Each time we call it, it sees only the context window: the full block of tokens we send in that request, up to a maximum size (its context length). Everything the model “knows” about our user, our product and the task at hand must be in that window.
Context engineering is the discipline of building that window well: choosing which pieces of information to include, in what form, in what order, and within what token budget, for every single model call. The term became popular in 2025, as people building AI agents noticed that most failures came from missing, wrong or cluttered context rather than from the model itself.
Think of it like briefing a brilliant new colleague A new colleague is smart but knows nothing about our company. Hand them a 500-page binder and they drown. Hand them nothing and they guess. Hand them a one-page brief with the goal, the three relevant documents, the customer's history and the tools they may use, and they do great work. Context engineering is writing that brief, automatically, for every call.
Our running example: Acme's support assistant answering “How do I fix the X200 leak, and does the warranty cover it?” We will build its context piece by piece.
The big picture: the context is assembled, not written
In a simple chat demo, the context is one prompt a person typed. In a real application, a program assembles the context from many sources at run time, and does so again for every call. For an agent, that can mean dozens of calls per task, each needing a fresh, well-chosen context.
Why context engineering matters
- The model cannot use what it cannot see. If the warranty document is not retrieved, the model either says it does not know or, worse, makes something up (a hallucination).
- More is not better. Irrelevant text distracts the model. Studies of long contexts show that models use information at the start and end of a long context more reliably than information buried in the middle (the “lost in the middle” effect), and that accuracy on many tasks tends to drop as the context grows, a problem people sometimes call context rot.
- Tokens cost money and time. Every token in the window is billed and adds latency. A bloated context can cost several times more than a lean one.
- Agents compound the problem. Each tool call adds output to the context. Without management, an agent's context fills with stale logs and old search results until it loses track of the goal.
Pause and think: Pause and predict: we double the number of retrieved documents from 5 to 10 “to be safe”. Can that make answers worse?
Yes. The extra 5 documents are usually less relevant. They push the useful ones toward the middle of the context, add distraction and contradictions, and increase cost and latency. More context helps only when the extra content is actually relevant.
Prompt engineering vs context engineering
Prompt engineering focuses on how we phrase instructions: the wording, examples and format of a prompt. Context engineering is broader: it is about the whole information environment of each call, most of which is not hand-written at all but retrieved, remembered or produced by tools.
The components of the context
| Component | What it is | In our support bot |
|---|---|---|
| System prompt | Standing instructions: role, rules, tone, output format | “Answer only from the docs below and cite doc ids.” |
| User message | The current request | “How do I fix the X200 leak…?” |
| Short-term memory | Recent conversation turns, possibly summarized | “user: the X200” |
| Long-term memory | Facts saved across sessions | “Prefers short answers. Plan: Pro.” |
| Retrieved knowledge | Documents found by search (RAG) | Gasket repair doc, warranty doc |
| Tool definitions | Names, descriptions and parameters of tools the model may call | lookup_order(order_id) |
| Tool results | Outputs returned by earlier tool calls | Order 1009: purchased 2025-08-03 |
| Output format | Schema or template for the answer | JSON with answer and sources |
Each component competes for the same token budget. Context engineering is largely the art of deciding how much room each gets for this particular call.
Code you can run: assembling context under a budget
This toy assembler builds the support bot's context. It keeps the system prompt and memory, trims history to the last two turns, retrieves documents by keyword overlap (a stand-in for real search), adds only relevant documents that fit the budget, and puts the question last. Token counting uses words as a rough stand-in for a tokenizer.
assemble_context.py
def tokens(text):
return len(text.split()) # rough stand-in for a real tokenizer
BUDGET = 70
system = "You are Acme's support assistant. Answer only from the docs below and cite doc ids."
memory = "User prefers short answers. Plan: Pro."
history = ["user: my blender leaks", "assistant: which model?", "user: the X200"]
docs = {
"D1": "An X200 leak is usually a worn gasket; fix it with the free gasket kit.",
"D2": "Our office is closed on public holidays.",
"D3": "X200 warranty: we cover parts for 2 years from purchase.",
"D4": "The X100 was discontinued in 2021 and is no longer covered.",
}
query = "how do I fix the X200 leak and does the warranty cover it"
STOP = {"how", "do", "i", "the", "and", "does", "it", "a", "is", "for", "with"}
def score(doc):
# Retrieval stand-in: count meaningful words shared with the query
words = lambda t: {w.strip(".,;:") for w in t.lower().split()} - STOP
return len(words(doc) & words(query))
ranked = sorted(docs, key=lambda d: score(docs[d]), reverse=True)
parts = [("system", system), ("memory", memory), ("history", " | ".join(history[-2:]))]
used = sum(tokens(t) for _, t in parts) + tokens(query)
for d in ranked: # add the best docs while they fit
cost = tokens(docs[d])
if score(docs[d]) >= 2 and used + cost <= BUDGET:
parts.append((f"doc {d}", docs[d]))
used += cost
parts.append(("question", query)) # question last, close to the answer
for name, text in parts:
print(f"[{name:8}] {tokens(text):2} tok | {text[:48]}")
print("scores:", {d: score(docs[d]) for d in ranked})
print(f"total: {used} of {BUDGET} tokens")Output:
[system ] 15 tok | You are Acme's support assistant. Answer only fr
[memory ] 6 tok | User prefers short answers. Plan: Pro.
[history ] 7 tok | assistant: which model? | user: the X200
[doc D1 ] 15 tok | An X200 leak is usually a worn gasket; fix it wi
[doc D3 ] 10 tok | X200 warranty: we cover parts for 2 years from p
[question] 13 tok | how do I fix the X200 leak and does the warranty
scores: {'D1': 3, 'D3': 3, 'D2': 0, 'D4': 0}
total: 66 of 70 tokensPause and think: D4 mentions “covered” and the question asks about coverage. Why did D4 score 0?
The query word is “cover” and D4 says “covered”; our toy retriever matches exact words only, and D4's other words (X100, discontinued) do not appear in the query. Real systems use embeddings or stemming, but the lesson is the same: selection decides what reaches the model, so we must test retrieval quality, not just the prompt.
Common patterns in context engineering
Four moves that cover most patterns
- Write (save outside the window): Store information outside the context so it is not lost: long-term memory, a scratchpad file, or an agent's to-do list. The context holds a pointer or a short note, not everything.
- Select (pull in what is relevant): Retrieve only the documents, memories and tools relevant to this call. RAG, memory search and choosing a subset of tools for the current step are all selection.
- Compress (keep the meaning, drop tokens): Summarize old conversation turns, trim verbose tool outputs to the needed fields, deduplicate repeated content. Context compaction is this move applied to long sessions.
- Isolate (split across contexts): Give sub-tasks to sub-agents with their own clean windows, and pass back only a short result. Keep big raw data (a 10 MB log) in a tool environment and send the model a summary.
- Just-in-time retrieval. Instead of preloading every file, give the agent tools (search, read file) and let it fetch what it needs when it needs it.
- Structured sections. Wrap each component in clear labels, for example
<docs>,<history>,<memory>, so the model can tell instructions from data. - Stable prefix first. Put unchanging parts at the start so prompt caching can reuse them.
Common mistakes and best practices
Common mistakes Stuffing everything into the window because the context limit is large. Retrieving many weakly relevant chunks. Letting raw tool outputs (full HTML pages, huge JSON) pile up. Keeping contradictory old instructions in history. Exposing 40 tools when the task needs 3, which confuses tool choice. Mixing untrusted data with instructions without labels, which also opens the door to prompt injection.
- Aim for the smallest set of high-signal tokens that lets the model do the task.
- Measure retrieval separately: is the right document in the context at all? Many “model errors” are retrieval errors.
- Budget each component (for example: system ≤ 1,500 tokens, docs ≤ 6,000, history ≤ 3,000) and enforce the budget in code.
- Put the question and key facts where the model attends best, usually near the end, and keep critical rules in the system prompt.
- Log the exact context of every call, so when an answer is wrong we can see what the model actually saw.
- Evaluate end to end with a test set of real questions whenever the assembly logic changes.
Real-world use Coding agents decide which files, error messages and test results to show the model at each step, and summarize the session when it gets long. Customer-support bots combine a customer profile from a database, the last few messages, policy documents from search and order data from tools. Research assistants hand sub-questions to sub-agents and merge their short reports.
Going one level deeper
Two skills turn the ideas above into daily practice: doing the budget arithmetic before we build, and reading a failed call to find which part of the context let us down.
A worked budget (illustrative numbers)
- Start from the window: Say our model has an 8,000-token window.
- Reserve the answer: Input and output share the window. We keep 1,000 tokens for the reply, which leaves 7,000 for input.
- Subtract the fixed parts: System prompt 600, tool definitions 900, output format 100. That is 1,600, so 5,400 tokens remain.
- Split the flexible parts: Question up to 200, memory 300, history 1,400, retrieved documents 3,500. Together: 5,400.
- Turn tokens into counts: With chunks of about 500 tokens, 3,500 tokens means at most 7 chunks. If retrieval returns 10, three must be dropped or compressed, and code should decide which.
The budget in our code example was only 70 tokens, but the logic is the same at any size: fixed parts first, then a hard cap for each flexible part.
| Symptom | Likely context cause | What to check in the log |
|---|---|---|
| A confident answer that is wrong | The needed document was never retrieved | Search the logged context for the fact. If it is absent, fix retrieval. |
| The fact is present but ignored | It sits in the middle of a long, cluttered window | Count the tokens around it. Cut weak chunks and move the fact near the question. |
| The model follows an outdated rule | An old instruction in the history contradicts the system prompt | Look for two rules that disagree. Summarize or drop the old one. |
| The answer stops mid-sentence | The input left too little room for the output | Compare the input token count with the window size. |
Practice: try it yourself
We will build a packer that works by priority. Each piece of context has a priority, a full version and, for some pieces, a short version. Under a tight budget the packer tries the full text first, then the short one, and only then drops the piece. At the end it lays out the kept pieces in a fixed order, with the question last.
practice_priority_packer.py
def tokens(text):
return len(text.split()) # rough stand-in for a tokenizer
# (name, priority, full text, short version). Lower number = more important.
items = [
("system", 0, "You are Acme's support assistant. Cite doc ids.", None),
("question", 0, "Does the warranty cover my X200 leak?", None),
("warranty", 1, "X200 warranty: we cover parts for 2 years from purchase. "
"A worn gasket counts as a part.", "X200 warranty: parts and gaskets, 2 years."),
("order", 1, "order_id=1009 status=delivered purchased=2025-08-03 carrier=FastShip "
"tracking=ZX81 gift_wrap=no", "Order 1009 purchased 2025-08-03."),
("history", 2, "user: my blender leaks | assistant: which model? | user: the X200",
"User has a leaking X200."),
("promo", 3, "This month all Acme mixers are cheaper for Pro members.", None),
]
ORDER = ["system", "promo", "history", "order", "warranty", "question"] # question last
def pack(budget):
chosen, used = {}, 0
for name, _, full, short in sorted(items, key=lambda it: it[1]):
for label, text in (("full", full), ("short", short)):
if text and used + tokens(text) <= budget: # try full, then short
chosen[name] = (label, tokens(text))
used += tokens(text)
break
return [n for n in ORDER if n in chosen], chosen, used
for budget in (70, 44, 30):
layout, chosen, used = pack(budget)
print(f"budget {budget}: used {used}")
print(" kept:", ", ".join(f"{n}({chosen[n][0]}, {chosen[n][1]})" for n in layout))
print(" dropped:", [it[0] for it in items if it[0] not in chosen])Output:
budget 70: used 60 kept: system(full, 8), promo(full, 10), history(full, 12), order(full, 6), warranty(full, 17), question(full, 7) dropped: [] budget 44: used 43 kept: system(full, 8), history(short, 5), order(full, 6), warranty(full, 17), question(full, 7) dropped: ['promo'] budget 30: used 28 kept: system(full, 8), order(full, 6), warranty(short, 7), question(full, 7) dropped: ['history', 'promo']
Now change it:
- Add
50to the list of budgets. Before running, predict which pieces are full, short or dropped. - Give
promopriority 0 and look at budget 30. Predict which piece gets squeezed out. Is that what we want from a wrongly set priority? - Give
promoa short version such as"Mixers cheaper for Pro."and look at budget 44 again. Predict whether it gets in.
Pause and think: At budget 30 the warranty document is kept in its short form, while the history is dropped. Why is that a better outcome for this question than full history and no warranty?
The question asks about warranty cover, so the warranty text is the one piece the model cannot answer without. The short version still carries the key fact: parts and gaskets, 2 years. The history only repeats that the user has a leaking X200, which the question already says. Compressing a high-value piece beats keeping a low-value piece whole.
Pause and think: The packer chooses pieces by priority but lays them out in the ORDER list. Why keep those two orders separate?
They answer different questions. Priority decides what survives when space is short. Layout decides where each survivor sits in the window: stable instructions first, so a cache can reuse them, and the question last, close to where the answer begins. If we laid the pieces out by priority, the question would sit near the top with all the documents after it.
Quick summary
The model's output can only be as good as its input. Context engineering treats the context window as a scarce, carefully managed resource: we write information to places outside it, select only what is relevant, compress what we keep, isolate sub-tasks into separate windows, and order and label everything clearly. Prompt wording still matters, but it is one piece of this larger system.
Key takeaways
- The model only knows what is in its context window for this call, so building that window is a core engineering task.
- Context is assembled at run time from system prompt, memory, history, retrieved docs, tools and tool results.
- More context is not better: aim for the smallest set of relevant, well-ordered tokens.
- Four core moves: write, select, compress and isolate.
- Prompt engineering is one part of context engineering; log and evaluate the exact context the model saw.
Key terms
- Context window: All the tokens an LLM can see in one call, up to its maximum context length.
- Context engineering: Designing what information goes into the context window for each model call, and in what form and order.
- Retrieval (RAG): Searching a knowledge source and inserting the most relevant results into the context.
- Long-term memory: Information saved outside the context across sessions and pulled back in when relevant.
- Lost in the middle: The tendency of models to use information in the middle of a long context less reliably than at its edges.
- Sub-agent isolation: Running a sub-task in a separate context window and returning only a compact result.
← 9.3 Prompt Caching: Reusing Computation Across API Calls · 9.5 Context Compaction: Fitting More Into a Finite Window →