Lesson 11.11 · 22 min
Multi-Agent Systems: Dividing Work Among Specialist Agents
If one AI agent is good, are five agents five times better, or five times the bill and five times the confusion?
In short: A multi-agent system is a group of LLM agents, each with its own role, context and tools, that communicate and coordinate to finish a task none of them handles as well alone. It shines when work splits into independent parts that can run in parallel or need separate expertise, but it costs more tokens, adds latency for coordination and creates new ways to fail. Start with one agent and add more only when you can name the reason.
The big picture
An AI agent is an LLM that runs in a loop: it decides on an action, calls a tool (search, code, an API), observes the result and repeats until the goal is met. One agent with good tools can do a lot. But as tasks grow, a single agent hits limits: its context window (the text it can hold at once) fills with notes from every subtask, it works through subtasks one at a time, and one prompt must make it an expert at everything.
Our running example: a market research assistant. A user asks, "Compare the top five electric-scooter makers on price, battery range, safety recalls and customer sentiment." That is four quite different research jobs, each needing many searches, plus a final report.
Think of it like a newsroom A newspaper does not have one reporter write the whole paper. An editor splits the work: one reporter covers prices, one covers safety, one interviews customers. They work at the same time, each with their own notebook. The editor reads their drafts and writes the final story. A multi-agent system is a newsroom of LLM agents, and the editor is usually an agent too.
What is a multi-agent system?
A multi-agent system (MAS) is a set of agents that work towards a shared goal by dividing work and exchanging information. The idea is old in computer science (robot swarms, trading agents), but here each agent is an LLM with its own prompt, tools and memory.
- Own context: each agent sees only what it needs, so its window stays focused.
- Own role: each gets a specialised system prompt ("You are a safety-recall researcher...").
- Own tools: the coder can run code, the researcher can search, the reviewer may have no tools at all.
- Shared goal: the outputs are combined into one result for the user.
Contrast this with a single agent with many tools: one loop, one context, one prompt. And with a fixed workflow (a pipeline where code, not the model, decides the order of LLM calls). Real systems often mix these: a fixed workflow whose steps are agents.
The three pillars
A useful way to describe any multi-agent system is by three pillars. If one is weak, the whole system wobbles.
| Pillar | Question it answers | In the scooter example |
|---|---|---|
| Agents | Who does the work, with what role, model and tools? | A lead planner, four researchers, a writer |
| Environment | What do they act on and share? | Web search, a notes store, the final report file |
| Interaction | How do they communicate and coordinate? | The lead sends tasks, researchers return summaries, the lead merges |
Interaction splits into two parts that people often blur. Communication is how information moves (messages, shared memory). Coordination is who decides what happens next (a boss, a fixed order, or a negotiation). We look at each below.
Common agent roles
| Role | Job | Typical tools |
|---|---|---|
| Orchestrator / planner | Breaks the goal into subtasks, assigns them, merges results | Spawn/assign agents, read results |
| Researcher | Gathers facts for one subtopic | Web search, document search |
| Executor / coder | Takes actions: writes and runs code, calls APIs | Code sandbox, APIs, file system |
| Critic / reviewer | Checks another agent's output for errors and gaps | Often none, or tests |
| Verifier | Checks facts or runs tests against a clear standard | Search, test runner |
| Writer / summariser | Turns collected material into the final answer | None |
Roles are not magic. A "critic" is the same kind of model with a prompt that says "find problems". It helps because a fresh context, without the author's chain of reasoning, often spots mistakes the author missed.
How agents communicate and coordinate
Communication styles (a later lesson goes deeper):
- Direct messages: agent A sends a message to agent B.
- Through a hub: every message goes via the orchestrator, which decides what to forward.
- Shared memory (blackboard): agents read and write a common store, such as a notes file or database.
- Broadcast: one agent sends the same message to all others.
Coordination patterns:
One run, step by step
- Plan: The lead reads the request and decides it splits into four independent research tracks.
- Delegate: Each subtask gets its own agent, prompt, tools and a budget (for example, at most 15 searches).
- Work in parallel: Agents search and read independently; none sees the others' raw notes.
- Return and merge: Each agent returns a compact result; the lead combines them into one table.
- Review and finish: A critic flags gaps; the lead re-delegates one fix if needed, then answers.
Multi-agent vs single agent: the trade-offs
The benefits are real but never free. The code below puts illustrative numbers on three effects: parallel speed-up, token overhead, and how fast the number of links grows when every agent may talk to every other agent. It also shows why voting only helps when agents make different mistakes.
tradeoffs.py
# Single agent vs multi-agent: time, cost, links and the value of independence.
import numpy as np
subtasks = [30, 25, 40, 20] # seconds each research subtask takes (illustrative)
plan, merge = 8, 10 # orchestrator overhead in seconds
tokens_per_task, overhead_tokens = 6000, 4000
single_time = sum(subtasks)
multi_time = plan + max(subtasks) + merge # workers run in parallel
single_tokens = sum(tokens_per_task for _ in subtasks)
multi_tokens = single_tokens + overhead_tokens * (len(subtasks) + 1)
print(f"time: single {single_time}s multi {multi_time}s")
print(f"tokens: single {single_tokens} multi {multi_tokens} "
f"(x{multi_tokens / single_tokens:.1f})")
for n in (3, 5, 10): # how many links must be managed?
print(f"{n} agents: peer-to-peer links {n * (n - 1) // 2:2d}, hub links {n - 1}")
# Three reviewer agents vote; each is right 70% of the time.
rng = np.random.default_rng(0)
trials, p = 100_000, 0.7
independent = rng.random((trials, 3)) < p
shared = rng.random(trials) < p # same model, same blind spot
correlated = np.where(rng.random((trials, 3)) < 0.8, shared[:, None], independent)
for name, votes in [("independent", independent), ("correlated", correlated)]:
majority = votes.sum(axis=1) >= 2
print(f"majority of 3, {name:11s}: {majority.mean():.3f}")Output:
time: single 115s multi 58s tokens: single 24000 multi 44000 (x1.8) 3 agents: peer-to-peer links 3, hub links 2 5 agents: peer-to-peer links 10, hub links 4 10 agents: peer-to-peer links 45, hub links 9 majority of 3, independent: 0.785 majority of 3, correlated : 0.707
Real systems show the same shape. Anthropic reported in 2025 that in its multi-agent research feature, multi-agent runs used roughly 15× the tokens of a normal chat, and that on a web-browsing benchmark, how many tokens were spent explained most of the variation in performance. In other words, part of the gain comes simply from spending more compute in parallel.
Pause and think: Using the code's numbers, why did multi-agent take 58 s and not 40 s (the longest subtask)?
Because the orchestrator adds overhead: 8 s to plan and 10 s to merge. 8 + 40 + 10 = 58. Coordination time is always added on top of the slowest worker.
Common mistakes
The biggest mistake: splitting work that is not independent If subtask B depends on details decided inside subtask A, giving them to separate agents means B works without those details. Two coding agents editing the same feature in parallel often make conflicting assumptions. Split only along clean seams, or pass the needed decisions explicitly.
- Too many agents, too early. Every extra role adds prompts, handoffs and cost. Many "teams" of five agents work better as one agent with good tools.
- Vague delegation. "Research scooters" leads to duplicated or missing work. Each subtask needs an objective, output format, tools and limits.
- Passing everything. Forwarding full transcripts between agents wastes context; forwarding too little loses key facts. Return condensed results with sources.
- No stopping rules. Agents that can call each other can loop forever. Set budgets for turns, tokens and time.
- Same model, same blind spots. Debate or voting among copies of one model gains little if they all make the same mistake (see the code output).
- No tracing. Without logs of every message and tool call, failures are almost impossible to diagnose.
When to use a multi-agent system
A good rule: start with one agent, measure where it fails, and add agents only for a reason you can name.
| Situation | Multi-agent? | Why |
|---|---|---|
| Many independent subtopics to research | Yes | Parallel, focused contexts |
| Task too big for one context window | Often | Each agent holds only its part |
| Parts need different permissions or tools | Yes | Isolation: the reader cannot write, the writer cannot browse |
| High-stakes output needing independent review | Yes, a reviewer | Fresh context catches errors |
| Tightly coupled task (one file, one design) | Usually no | Shared context matters more than parallelism |
| Simple Q&A or low budget | No | Overhead outweighs the gain |
Where you see it in practice Deep-research features in AI assistants use a lead agent that spawns parallel search agents. Coding assistants delegate exploration of a large codebase to helper agents and keep the main conversation clean. Customer-service platforms hand a conversation from a triage agent to a billing or technical specialist.
Pause and think: A team wants three agents to co-write one 200-line function: one for the first part, one for the middle, one for the end. Good idea?
Probably not. The parts are tightly coupled (shared variables, shared assumptions), so separate contexts will drift apart. One agent, perhaps with a separate reviewer agent, is the better design.
Worked example, step by step
The mistakes list says “vague delegation leads to duplicated or missing work”. Let us watch that happen with small made-up numbers, and then fix it. The lead agent in our scooter example sends two researchers out with one-line tasks: “research prices” and “research battery range”.
| Maker | Price researcher | Range researcher |
|---|---|---|
| Zip | 400 | 30 |
| Volt | 599 | 28 |
| Glide | 430 | 38 |
The table looks complete, so the lead writes: “Volt costs the most and has the shortest range.” Both claims are wrong, and no agent made a search error.
Finding the hidden disagreement
- Different units: The range researcher reported Zip and Glide in kilometres but copied Volt's figure in miles from a US page. 28 miles is about 45 km, the longest range of the three, not the shortest.
- Different currencies: The price researcher mixed euros and dollars in the same way. Volt's 599 is in dollars; the others are in euros.
- Different product lists: The price researcher looked at each maker's cheapest model; the range researcher looked at each maker's best-selling model. The two columns describe different scooters.
- Why nobody noticed: Each researcher worked in its own context and returned bare numbers. The decisions that mattered (which model, which unit) stayed inside each worker and were lost at the handoff.
The fix is not a smarter model. It is a shared contract written by the lead and sent in every brief, so that separate contexts still make the same choices.
Practice: try it yourself
The earlier code measured trade-offs with arithmetic. Here we run the supervisor pattern itself, in miniature: a lead with a plan, two researcher agents with private notes, a critic with a fresh view, and a budget of one re-delegation. The agents are scripted functions; the coordination around them is real.
practice_supervisor.py
# Supervisor pattern: plan -> delegate -> merge -> critic -> one re-delegation.
MAKERS = ["Zip", "Volt", "Glide"]
WEB = {("price", "Zip"): 400, ("price", "Volt"): 550, ("price", "Glide"): 480,
("range_km", "Zip"): 30, ("range_km", "Volt"): 45}
DEEP_WEB = {("range_km", "Glide"): 38} # only found with a deeper search
def researcher(topic, makers, deep=False):
"""A worker agent. Its notes are its private context; it returns a summary."""
notes, found = [], {}
for maker in makers:
value = WEB.get((topic, maker))
if value is None and deep:
value = DEEP_WEB.get((topic, maker))
notes.append(f"searched {topic} for {maker}: {value}") # never leaves the worker
if value is not None:
found[maker] = value
return found, len(notes)
def critic(table):
"""A reviewer with a fresh context: lists every missing cell."""
return [(topic, m) for topic in table for m in MAKERS if m not in table[topic]]
messages, table = 0, {}
plan = ["price", "range_km"] # the lead agent's plan (scripted)
for topic in plan: # independent, so they could run in parallel
table[topic], searches = researcher(topic, MAKERS)
messages += 2 # one task out, one result back
print(f"researcher({topic}): {searches} private notes -> returned {table[topic]}")
gaps = critic(table)
messages += 2
print("critic found gaps:", gaps)
for topic, maker in gaps[:1]: # budget: at most one re-delegation
fix, _ = researcher(topic, [maker], deep=True)
table[topic].update(fix)
messages += 2
print(f"re-delegated {topic}/{maker} -> {fix}")
print("final table:", table)
print("gaps left:", critic(table), "| messages through the lead:", messages)Output:
researcher(price): 3 private notes -> returned {'Zip': 400, 'Volt': 550, 'Glide': 480}
researcher(range_km): 3 private notes -> returned {'Zip': 30, 'Volt': 45}
critic found gaps: [('range_km', 'Glide')]
re-delegated range_km/Glide -> {'Glide': 38}
final table: {'price': {'Zip': 400, 'Volt': 550, 'Glide': 480}, 'range_km': {'Zip': 30, 'Volt': 45, 'Glide': 38}}
gaps left: [] | messages through the lead: 8Now change it:
- Delete the
("price", "Volt"): 550entry fromWEB. Predict what the critic finds, which gap gets re-delegated, and what “gaps left” shows at the end. - Change
gaps[:1]togaps[:0](no re-delegation budget). Predict the final table and the message count. - Add
"recalls"toplan, with no recall data inWEB. Predict how many gaps the critic reports and how many are left at the end. What does this say about a critic paired with a tight budget?
Pause and think: Each researcher wrote 3 private notes, but the lead only ever saw a small dictionary. What is gained by that, and what could be lost?
Gained: the lead's context stays small and focused, however much searching each worker did. With real agents those notes could be thousands of tokens of search results. Lost: anything the worker saw but did not put in its result, such as “the Glide page showed two different ranges”. That is why results should carry sources and open questions, and why the worker's full notes should still be logged for debugging.
Pause and think: The critic found the missing Glide range, but it could not have found a wrong Volt price. Why not, and what would a critic need in order to catch that?
This critic only checks that every cell is filled. A wrong number fills its cell just as well as a right one. To catch wrong values the critic needs evidence to check against: the sources behind each number, a second independent lookup, or rules such as “a price must be in the stated currency”. A reviewer is only as strong as the standard it checks against.
Quick summary
A multi-agent system is several LLM agents with their own roles, contexts and tools, working towards one goal. Think in three pillars: agents, environment, interaction (communication + coordination). Common roles are orchestrator, researcher, executor, critic, verifier and writer. Communication can be direct, through a hub, broadcast or via shared memory; coordination is usually a central supervisor. Multi-agent buys speed on wide tasks, focused contexts and specialisation, and pays in tokens, latency for coordination and new failure modes. Use it for parallel, separable work, not for tightly coupled tasks.
Key takeaways
- A multi-agent system = several agents with their own roles, contexts and tools, plus communication and coordination.
- Three pillars: agents, environment, interaction (communication + coordination).
- Centralised supervisor is the default coordination pattern; others add flexibility and risk.
- Gains: parallel speed on wide tasks, focused contexts, specialisation. Costs: tokens, overhead, new failure modes.
- Split only along independent seams; tightly coupled work belongs in one agent.
- Start with one agent and add more only for a reason you can name and measure.
Key terms
- Multi-agent system (MAS): Several AI agents that divide work and exchange information to reach a shared goal.
- Orchestrator: The agent (or code) that plans, assigns subtasks and merges results.
- Communication: How information moves between agents: messages, a hub, broadcast or shared memory.
- Coordination: How the system decides who does what and when: supervisor, hierarchy, peers or voting.
- Blackboard (shared memory): A common store that all agents can read and write.
- Handoff: Passing control of a task or conversation from one agent to another.
← 11.10 Open Knowledge Format: Structured Agent-to-Agent Communication · 11.12 Subagents: Delegating Tasks Within an Agent Network →