Modern AI Engineering

Lesson 11.12 · 20 min

Subagents: Delegating Tasks Within an Agent Network

Why would an AI agent hand part of its own job to another AI agent, and then throw away almost everything that helper read?

In short: A subagent is a helper agent that a main agent starts for one focused subtask. It runs with its own fresh context window, its own instructions and often a restricted set of tools, and it returns only a short result. This keeps the main agent's context clean, allows parallel work and lets us give each helper only the permissions it needs.

What is an AI agent?

An AI agent is a large language model (LLM) placed in a loop with tools. Each turn, the model reads its context, decides on an action (for example "search the code for login"), a tool runs that action, and the tool's output is appended to the context. The loop repeats until the model decides the goal is done.

That last point is the root of this lesson. Everything the agent reads stays in its context window, the fixed number of tokens the model can attend to at once. A big file read, a long search result or a noisy test log all take space, and they stay there for the rest of the task.

What are AI subagents?

A subagent is an agent that another agent (the main agent, also called the parent or orchestrator) creates to do one focused subtask. The subagent gets a task description, works through its own loop with its own tools, and returns a result. Then it ends.

  • Fresh context: the subagent starts with an empty window plus its instructions. It does not see the whole parent conversation.
  • Own instructions: a system prompt tailored to its job, such as "You are a code explorer. Report file paths and line numbers only."
  • Own tools: often fewer than the parent, for example read-only file access.
  • Short result: only the final answer goes back to the parent. Its working notes are discarded.

Think of it like sending an intern to the archive A lawyer preparing a case does not personally read 3,000 pages of old records. She sends an intern: "Find every contract with Acme from 2019 and tell me the renewal dates." The intern spends a day in the archive and comes back with one page. The lawyer's desk stays clear, and she keeps thinking about the case. The intern is the subagent; the one page is the returned result.

Why do we need subagents?

Our running example: a coding assistant working in a large repository. The user asks, "Add rate limiting to the login endpoint." Before changing anything, the assistant must find where login is handled, how other endpoints do rate limiting, and which tests cover login.

  • Context pollution: exploring means reading many files. If all of that lands in the main context, the window fills with irrelevant code, cost rises on every later turn, and quality can drop because the important parts are buried.
  • Focus: a subagent with one narrow goal and a clean context is less distracted than a main agent juggling the whole conversation.
  • Parallelism: three independent questions can go to three subagents at the same time.
  • Least privilege: an explorer subagent can be given read-only tools, so it cannot edit anything by mistake.
  • Specialisation and cost: a subagent can use a different prompt or even a smaller, cheaper model for simple jobs like searching.

Pause and think: The main agent reads 40 files (about 80,000 tokens) to find one function. Without subagents, what happens to every later turn of the conversation?

All 80,000 tokens stay in the context, so every later model call re-reads them: it costs more, is slower, and the useful information is diluted. A subagent would have read them in its own window and returned just "the function is in auth/login.py, line 42".

How do subagents work?

Under the hood, a subagent is usually exposed to the main agent as a tool. The main model calls something like spawn_subagent(task="...", type="explorer"). The framework then runs a whole separate agent loop and returns that agent's final message as the tool result.

The life of a subagent

  1. Decide to delegate: The main agent sees a subtask that is self-contained and would produce lots of noise, e.g. "find where login is handled".
  2. Write the brief: It writes a clear task: goal, what to return, limits. The brief is all the subagent knows, so it must include any needed context.
  3. Spawn with a fresh context: The framework starts a new loop: subagent system prompt + the brief, plus its allowed tools.
  4. Work independently: The subagent searches, reads and reasons over many turns. All of that stays in its own window.
  5. Return a summary: It ends with a compact answer, ideally structured (paths, line numbers, findings, open questions).
  6. Continue the main task: The main agent receives the result as a tool output and keeps going, with a context that grew by only a few hundred tokens.

Many tools make this concrete. In Claude Code, for example, we can define a custom subagent as a Markdown file with YAML frontmatter (a name, a description that says when to use it, an optional tool list and model) in a .claude/agents/ folder, and the main agent delegates to it when a task matches. Other frameworks expose the same idea as "agents as tools". The exact configuration differs by product.

Example use case, in code

Let us measure the context effect. We build a fake repository of 30 files × 400 lines, run three search tasks, and compare the main agent's context when tool output goes straight into it versus when a subagent does the reading and returns a one-line summary.

subagent_context.py

# Main agent context with and without subagents (stdlib only, fake codebase).
import random
random.seed(7)
REPO = {f"src/module_{i}.py": [f"line {j}: " + ("def login" if random.random() < 0.01 else "x = 1")
for j in range(400)] for i in range(30)}
def tokens(lines):                       # rough rule: ~0.75 words per token
return round(sum(len(l.split()) for l in lines) / 0.75)
def explore(keyword):
"""Read every file, find the keyword. Returns (raw tool output, short summary)."""
raw, hits = [], []
for path, lines in REPO.items():
raw += lines                      # everything the explorer had to read
hits += [f"{path}:{n}" for n, l in enumerate(lines) if keyword in l]
summary = [f"'{keyword}' found {len(hits)} times, e.g. {hits[:2]}"]
return raw, summary
tasks = ["def login", "line 399", "def logout"]
main_without, main_with, sub_peak = 2000, 2000, 0    # 2,000 tokens of user chat
for t in tasks:
raw, summary = explore(t)
main_without += tokens(raw)                       # tool output lands in main context
main_with += tokens(summary)                      # only the summary comes back
sub_peak = max(sub_peak, tokens(raw))             # the subagent's own window
print(summary[0][:70])
print(f"\nmain context WITHOUT subagents: {main_without:,} tokens")
print(f"main context WITH subagents:    {main_with:,} tokens")
print(f"largest single subagent context: {sub_peak:,} tokens (then discarded)")

Output:

'def login' found 128 times, e.g. ['src/module_0.py:106', 'src/module_
'line 399' found 30 times, e.g. ['src/module_0.py:399', 'src/module_1.
'def logout' found 0 times, e.g. []
main context WITHOUT subagents: 241,487 tokens
main context WITH subagents:    2,031 tokens
largest single subagent context: 79,829 tokens (then discarded)

Notice what the numbers do not say. Total tokens processed across all agents are about the same (or slightly higher, because each subagent also reads its own prompt). Subagents save the main context, not total compute. In this toy example, without subagents the main context (241k tokens) would overflow many models' windows; with them, it stays tiny.

Benefits and challenges

The handoff is the weak point A subagent knows only what is in its brief. If the main agent writes "check the auth code" without saying what to look for, the subagent may return a confident but useless summary. And because the main agent never sees the raw files, it cannot easily notice what was missed. Write briefs like you would for a capable new colleague with zero background.

  • Lost context: the subagent cannot see decisions made earlier in the main conversation unless they are in the brief.
  • Summary errors: a wrong or incomplete summary is trusted as fact by the main agent.
  • Conflicting edits: two subagents that both write to the same files in parallel can overwrite each other.
  • Cost and latency: each spawn re-sends a system prompt and tool definitions and adds turns.
  • Debugging: failures are spread across several traces; we need logs of every subagent run.
  • Runaway delegation: subagents that spawn subagents can explode in cost; many tools limit or forbid nesting.

Best practices

  • One clear job per subagent, with a name and description that make it obvious when to use it.
  • Write complete briefs: goal, relevant background, what not to do, and the exact shape of the answer.
  • Ask for structured results: file paths, line numbers, short findings, confidence and open questions, so the main agent can verify.
  • Give the fewest tools needed: read-only for explorers and reviewers; write access only where edits are the point.
  • Prefer read-heavy tasks for parallel subagents; keep writes in the main agent or serialise them.
  • Set budgets: maximum turns, tokens or time per subagent.
  • Log everything: keep each subagent's transcript for debugging even though the main context discards it.
  • Do not over-delegate: a two-second lookup is cheaper done directly.

Here is what a good brief looks like for our rate-limiting example. Notice that it carries the background the subagent cannot see, says what not to do, and fixes the shape of the answer so the main agent can check it quickly.

A brief for an explorer subagent (illustrative)

Goal: find how HTTP rate limiting is currently implemented in this repo.
Background: we will add rate limiting to POST /login next. Python 3.12, FastAPI.
Do: search for middleware, decorators or Redis counters used for limits.
Do not: edit any file, or read the frontend/ folder.
Return (max 15 lines):
- file:line of each existing limiter and what it limits
- how limits are configured (env vars, settings file)
- anything that looks broken or unused
- your confidence (high / medium / low) and open questions

Compare that with "look at rate limiting". The vague version forces the subagent to guess the purpose, which files matter and when to stop, and the main agent then receives an answer it cannot easily verify. A good brief usually costs a hundred tokens and saves thousands.

Where subagents are used today Coding assistants spawn explorer subagents to map a codebase and reviewer subagents to check a diff with a fresh eye. Deep-research products send parallel search subagents out on different subtopics. Data agents delegate "summarise this 500-page PDF" to a helper so the main analysis stays focused.

Pause and think: Two subagents are asked to refactor two different functions in the same file at the same time. What can go wrong, and how would we avoid it?

Their edits can conflict: one writes the file based on an old version and overwrites the other's change. Avoid it by giving parallel subagents read-only work, by assigning non-overlapping files, or by running the editing subagents one after another.

Worked example, step by step

The best-practice list ends with “do not over-delegate”. How do we know when a task is too small to hand off? We can estimate it with simple arithmetic. All numbers below are illustrative, and the model is deliberately rough.

The idea: anything that lands in the main context is re-read on every later turn. So the cost of doing a task directly is its raw output multiplied by the turns still to come. The cost of delegating is a one-off: the subagent's start-up plus its reading, and then only the short summary is re-read.

We take S = 1,600, A = 150 and T = 10 turns left, and compare two tasks.

Two tasks, same formula (illustrative numbers)
TaskRDirect: R × TDelegated: S + R + A × TBetter choice
Explore 20 files to find the login flow40,000400,0001,600 + 40,000 + 1,500 = 43,100Delegate
Read one short config file3003,0001,600 + 300 + 1,500 = 3,400Do it directly

Reading the result

  1. Big raw output, many turns left: delegate: The exploration would be re-read ten times in the main context. Delegating cuts the estimate by roughly nine tenths.
  2. Small raw output: do it directly: For the config file, the start-up cost alone is more than five times the file. Delegation costs more and adds a handoff that can lose detail.
  3. Few turns left changes the answer: With T = 1, the exploration costs 40,000 directly and 41,750 delegated. Near the end of a task, even a big read may not be worth a subagent.
  4. Tokens are not the only reason: The estimate ignores the other benefits: parallel work, a fresh view, and restricted tools. A read-only reviewer can be worth spawning even when the token sums are equal.

Practice: try it yourself

The earlier code measured context sizes. Now we build the mechanism: a spawn_subagent function that the main agent calls like a tool. The subagent starts from a fresh context holding only the brief, runs its own loop with an allow-list of tools, and hands back one line. Its model is scripted, and we make it overstep once, so we can see the tool restriction work.

practice_spawn_subagent.py

# A subagent exposed as a tool: fresh context, restricted tools, short result.
REPO = {"auth/login.py": "def login(user): return check_password(user)",
"auth/limits.py": "def rate_limit(key, per_minute): ...",
"tests/test_login.py": "def test_login(): assert login('ann')"}
ORIGINAL = dict(REPO)
def read_file(path):
return REPO[path]
def write_file(path, text):
REPO[path] = text
return "written"
TOOLS = {"read_file": read_file, "write_file": write_file}
def spawn_subagent(brief, allowed):
"""Run a separate agent loop. Only the summary goes back to the caller."""
context = [brief]                            # fresh: it sees the brief, nothing else
script = [("read_file", ("auth/login.py",)), ("read_file", ("auth/limits.py",)),
("write_file", ("auth/login.py", "# tidy up")),     # it oversteps
("read_file", ("tests/test_login.py",))]            # scripted model choices
seen = {}
for tool, args in script:                    # the subagent's own loop
if tool not in allowed:
context.append(f"DENIED: {tool} is not in your tool list")
continue
seen[args[0]] = TOOLS[tool](*args)
context.append(f"{args[0]}: {seen[args[0]]}")
handler = [p for p, text in seen.items() if "def login" in text]
limiter = [p for p, text in seen.items() if "rate_limit" in text]
return f"login handler: {handler}; existing limiter: {limiter}", context
main_context = ["user: add rate limiting to the login endpoint"]
brief = "Find the login handler and any existing rate limiter. Do not edit files."
summary, sub_context = spawn_subagent(brief, allowed={"read_file"})
main_context.append("subagent result: " + summary)   # arrives like a tool result
size = lambda ctx: sum(len(item.split()) for item in ctx)
print("summary:", summary)
print("subagent context:", len(sub_context), "items,", size(sub_context), "words (discarded)")
print("main context:    ", len(main_context), "items,", size(main_context), "words")
print("denied calls:", sum("DENIED" in c for c in sub_context), "| repo unchanged:", REPO == ORIGINAL)

Output:

summary: login handler: ['auth/login.py']; existing limiter: ['auth/limits.py']
subagent context: 5 items, 36 words (discarded)
main context:     2 items, 16 words
denied calls: 1 | repo unchanged: True

Now change it:

  • Spawn with allowed={"read_file", "write_file"}. Predict the last output line and what the summary now says about the login handler. Why did one extra permission make the answer worse, not just the repo?
  • Remove the tests/test_login.py step from script and add a third finding to the summary that lists files containing test_. Predict what it reports. What does the main agent learn about tests in that case?
  • Append sub_context to main_context instead of the summary (use main_context += sub_context). Predict the new size of the main context in items and words. Which benefit of subagents did we just give up?

Pause and think: The brief said “Do not edit files”, and the tool list also blocked write_file. The subagent still tried to write. Which of the two protections actually stopped it, and what is the brief's instruction still good for?

The allow-list stopped it: the loop refused the call, so the repo stayed unchanged whatever the model wanted. The sentence in the brief is guidance, which a model usually follows but can ignore. It is still useful: it tells the subagent not to waste turns trying, and it documents the intent. For anything that must not happen, the restriction has to live in the tools, not only in the words.

Pause and think: The main agent sees a 2-item context and a confident one-line summary. From that alone, can it tell that the subagent was denied a write, or which files it never opened?

No. Both facts live only in the subagent's context, which is discarded. The summary mentions neither. This is the handoff risk in practice: the main agent trusts a result it cannot inspect. Two habits reduce it: ask in the brief for a structured result that lists what was checked and what was not, and keep the subagent's full transcript in a log so a person can look when something seems off.

Key takeaways

  • A subagent is a helper agent spawned for one focused subtask, with its own fresh context, prompt and tools.
  • Only its short result returns to the main agent, keeping the main context clean.
  • Subagents enable parallel work and least-privilege tool access; they do not reduce total compute.
  • The brief is everything the subagent knows: make it complete and ask for structured answers.
  • Avoid parallel writes to the same files, set budgets and log every subagent run.

Key terms

  • Subagent: An agent started by another agent to do one subtask in its own context and return a result.
  • Main agent: The agent that holds the user conversation, plans the work and delegates to subagents.
  • Context window: The maximum number of tokens a model can read in one call.
  • Context pollution: Filling the context with large, mostly irrelevant tool output, which raises cost and can lower quality.
  • Brief: The task description handed to a subagent; its only knowledge of the larger job.
  • Least privilege: Giving each agent only the tools and permissions its job requires.

← 11.11 Multi-Agent Systems: Dividing Work Among Specialist Agents · 11.13 Agent Communication: Protocols and Message Formats →