Lesson 11.14 · 23 min
AI Orchestration: Coordinating Agents, Tools, and Flows
A real AI product is rarely one model call. Who decides which call happens first, which ones run together, and what happens when one of them fails?
In short: AI orchestration is the coordination layer that connects models, tools, data, memory and agents into one reliable workflow. It decides the order of steps, passes data between them, handles branching, retries and errors, and keeps state. Five patterns cover most systems: sequential, parallel, conditional, loop and orchestrator-worker.
What is AI orchestration?
AI orchestration is the coordination of several AI components (LLM calls, tools, retrieval, memory, other agents) so that together they complete a task. The orchestrator is the part that holds the plan: it decides what runs, in what order, with which inputs, and what to do with the outputs.
Our running example: a customer-email assistant for an online shop. For each incoming email it must detect the language, classify the request, look up the order, draft a reply, check the reply for policy problems, and either send it or pass it to a human. That is six or more steps, some with tools, some with conditions.
Think of it like a conductor and an orchestra Each musician (a model, a tool, a database) is skilled at one thing. The conductor does not play an instrument. They decide when each section comes in, how loud, and how the parts fit. Without a conductor, skilled musicians still produce noise. Orchestration is the conductor for AI components.
Why do we need AI orchestration?
- One call is not enough: real tasks need retrieval, tools, checks and formatting, not just one prompt.
- Reliability: models fail, time out and sometimes return malformed output. Someone must retry, fall back or stop.
- Data flow: the order number found in step 2 must reach the lookup tool in step 3.
- Cost and speed: independent steps should run in parallel; cheap models should handle easy steps.
- State and memory: long tasks must remember what already happened, and resume after a crash.
- Observability and control: we need logs, human approval points and limits on spending.
Without an orchestration layer, this logic ends up scattered across ad-hoc scripts, which is hard to test, change or debug.
AI orchestration vs AI agents
These words overlap, so let us be precise. An AI agent is an LLM that decides its own next action in a loop, using tools until it reaches a goal. Orchestration is the broader job of coordinating components. The orchestrator can be code (a fixed workflow written by developers) or an LLM (an agent that plans and delegates at run time).
Components of AI orchestration
| Component | Role | In the email assistant |
|---|---|---|
| Models | Do the reasoning and generation in each step | A small model to classify, a larger one to draft |
| Tools / integrations | Act on the world or fetch data | Order database, email API |
| Data / retrieval | Bring in knowledge | Return policy documents |
| Memory / state | Remember what happened so far | Email, language, order, draft, check results |
| Control logic | Order, branches, loops, parallelism | "If refund request, go to refund branch" |
| Error handling | Retries, timeouts, fallbacks | Retry the order lookup twice, then escalate |
| Observability | Logs, traces, metrics, costs | A trace of every step per email |
| Guardrails / human-in-the-loop | Checks and approvals | Human approves refunds above $200 |
How AI orchestration works
One email through the orchestrator
- Receive and initialise state: The orchestrator creates a state object:
{email, customer_id}and a trace id. - Plan or follow the graph: In a workflow, the next node is fixed by code. In an agent, an LLM chooses the next action.
- Run a step: Call a model or tool with inputs taken from the state, e.g. classify the email as
refund. - Update state: Write the result back:
state.intent = refund. - Decide what is next: Branch on the state (refund branch), run independent steps in parallel, or loop back for a revision.
- Handle failure: On error or timeout: retry, use a fallback model, or route to a human.
- Finish: Return the result, store the trace, record cost and latency.
Patterns of AI orchestration
Almost every orchestrated system is built from five patterns. They nest: a loop can contain a parallel step; an orchestrator-worker can run inside one branch.
patterns.py
# Five orchestration patterns with fake "LLM steps" (stdlib only).
# Each step returns (result, seconds); we add up simulated time instead of sleeping.
def step(name, secs):
return lambda x: (f"{name}({x})", secs)
def sequential(steps, x): # output of one step feeds the next
total = 0
for s in steps:
x, t = s(x); total += t
return x, total
def parallel(steps, x): # same input, run at once, wait for slowest
outs = [s(x) for s in steps]
return [o for o, _ in outs], max(t for _, t in outs)
def conditional(x): # a router picks one branch
branch = step("refund_agent", 2) if "refund" in x else step("faq_agent", 1)
return branch(x)
def loop(draft, max_iters=5, target=0.8): # refine until a checker is satisfied
score, i = 0.5, 0
while score < target and i < max_iters:
score, i = round(score + 0.12, 2), i + 1 # each revision helps a bit
return f"{draft} after {i} revisions (score {score})", i * 3
def orchestrator_worker(question): # plan subtasks at run time, then merge
subtasks = [w for w in ["price", "range", "recalls"] if w in question]
results, t = parallel([step(f"worker:{s}", 4) for s in subtasks], question)
return f"merge({len(results)} results)", 2 + t + 2 # plan + work + merge
print("sequential :", sequential([step("extract", 2), step("summarise", 3), step("translate", 2)], "doc"))
print("parallel :", parallel([step("sentiment", 2), step("topics", 3), step("pii_check", 1)], "review"))
print("conditional:", conditional("I want a refund"), conditional("opening hours?"))
print("loop :", loop("essay"))
print("orch-worker:", orchestrator_worker("compare price and recalls"))Output:
sequential : ('translate(summarise(extract(doc)))', 7)
parallel : (['sentiment(review)', 'topics(review)', 'pii_check(review)'], 3)
conditional: ('refund_agent(I want a refund)', 2) ('faq_agent(opening hours?)', 1)
loop : ('essay after 3 revisions (score 0.86)', 9)
orch-worker: ('merge(2 results)', 8)Pause and think: Three independent steps take 4 s, 6 s and 5 s. How long do they take sequentially and in parallel (ignoring overhead)?
Sequentially 4 + 6 + 5 = 15 s. In parallel max(4, 6, 5) = 6 s. Parallel only works because the steps do not need each other's outputs.
Tools for AI orchestration
We can orchestrate with plain code, and for small systems that is often best. As systems grow, frameworks help with state, retries, tracing and human approval. The landscape changes fast; these are examples of categories, not a ranking.
| Category | Examples | What they give you |
|---|---|---|
| Graph / workflow frameworks for LLMs | LangGraph, LlamaIndex Workflows | Steps as nodes, explicit state, branches and loops, checkpoints |
| Multi-agent frameworks | CrewAI, Microsoft AutoGen / Agent Framework | Roles, agent conversations, delegation |
| Vendor agent SDKs | OpenAI Agents SDK, Claude Agent SDK, Google ADK | Agent loops, tools, handoffs or subagents, tracing |
| Durable workflow engines | Temporal, Airflow, Prefect | Retries, scheduling, resuming long-running jobs after crashes |
| Plain code | Python functions + asyncio | Full control, no dependencies; you build logging and retries |
Challenges in AI orchestration
- Error propagation: a wrong classification early on sends everything after it down the wrong path.
- Latency stacking: each sequential LLM call adds seconds; long chains feel slow.
- Cost growth: loops, retries and orchestrator calls multiply tokens.
- State management: keeping the right data available to each step without passing huge contexts around.
- Non-determinism: the same input can take different paths, which makes testing harder.
- Debugging: failures hide across many steps unless every step is traced.
- Over-engineering: a five-agent graph for a task one prompt can do.
The common mistake: loops and agents without limits A loop that waits for "good enough" and an agent that decides when it is done can both run forever, especially when the checker is strict and the generator cannot satisfy it. Always set maximum iterations, timeouts and a token or cost budget, and define what happens when a limit is hit (usually: hand off to a human).
Best practices
- Start with the simplest thing: one prompt, then a chain, then add patterns as measured needs appear.
- Prefer code for control when the path is known; use an LLM orchestrator only where flexibility pays.
- Make state explicit: a typed state object that every step reads and writes.
- Validate between steps: check output formats before passing them on.
- Parallelise independent work, and handle partial failures in the merge.
- Bound every loop: iterations, time and cost.
- Trace everything: inputs, outputs, latency and cost per step, with one trace id per request.
- Put humans at the risky points: approvals for money, deletions or external messages.
- Evaluate the whole pipeline, not only single steps: a test set of real inputs with expected outcomes.
Orchestration in the wild Retrieval-augmented chatbots chain retrieve → rerank → generate → check. Coding agents loop write → run tests → fix. Document pipelines fan out pages to parallel extractors and merge the fields. Support systems route by intent to specialist agents with different tools.
Going one level deeper
The components table lists “retries, timeouts, fallbacks” in a single row. They deserve a closer look, because they decide whether a chain of steps is dependable. Let us work it out with small numbers. These are simple probability sums, not measurements.
Take a sequential chain of 3 steps. Each step works 90% of the time, and failures are passing glitches such as a timeout, so one attempt failing says nothing about the next. The chain needs all three steps, so it succeeds 0.9 × 0.9 × 0.9 ≈ 73% of the time.
Now let each step try again once when it fails. A step fails only if both attempts fail: 0.1 × 0.1 = 0.01. So each step now works 99% of the time, and the chain works 0.99³ ≈ 97% of the time.
That is a large gain for a few lines of code. But the sum hides three conditions, and each one is a common source of bugs.
What retries cannot do
- They only fix passing failures: If a step fails because its input is wrong, such as an order number that does not exist, it will fail the same way every time. Retrying wastes time and money. This needs a different path: a fallback or a human.
- They are only safe for repeatable steps: Looking up an order twice is harmless. Sending an email or issuing a refund twice is not. Before retrying a step that changes something, check whether the first attempt went through.
- They cost time: If one attempt may take 5 s before timing out, a step with 2 retries may take 15 s in the worst case. Three such steps in a row could take 45 s. Set a time budget for the whole request, not only per step.
- They need a last resort: When the retries run out, the orchestrator must still do something sensible: use a simpler method, return a partial answer, or hand off to a person. That last branch is the fallback.
Practice: try it yourself
The earlier code compared patterns by their timing. Here we build a tiny workflow engine for the email assistant, with the parts that timing sums leave out: a state object every step reads and writes, a router, a retry around a flaky tool, a fallback to a human, and a trace of what happened. The model steps are scripted functions.
practice_email_workflow.py
# A small workflow engine: explicit state, routing, retries, a fallback and a trace.
ORDERS = {"1042": 250, "2001": 80}
attempts = {"n": 0}
def classify(state):
return {"intent": "refund" if "refund" in state["email"] else "other"}
def lookup_order(state): # a flaky tool: calls 1, 3, 5... time out
attempts["n"] += 1
if attempts["n"] % 2 == 1:
raise TimeoutError("order DB timeout")
order_id = next(w for w in state["email"].split() if w in ORDERS)
return {"amount": ORDERS[order_id]}
def draft_refund(state): return {"reply": f"Your refund of {state['amount']} is on its way."}
def faq(state): return {"reply": "We are open 9 to 5, Monday to Friday."}
def escalate(state): return {"reply": "HANDED TO A HUMAN"}
def run_step(fn, state, trace, retries=1):
"""Run one step, write its result into the state, retry on failure."""
for attempt in range(1, retries + 2):
try:
state.update(fn(state))
trace.append(f"{fn.__name__} ok" + (f" (try {attempt})" if attempt > 1 else ""))
return True
except Exception as error:
trace.append(f"{fn.__name__} failed: {error}")
return False
def handle(email, retries=1):
state, trace = {"email": email}, []
run_step(classify, state, trace)
if state["intent"] != "refund": # conditional
run_step(faq, state, trace)
elif run_step(lookup_order, state, trace, retries) and state["amount"] <= 200:
run_step(draft_refund, state, trace)
else: # fallback: a person
run_step(escalate, state, trace)
return state["reply"], trace
for email in ["refund for order 2001 please", "refund order 1042", "what are your hours?"]:
reply, trace = handle(email)
print(f"{reply}\n trace: {' -> '.join(trace)}")Output:
Your refund of 80 is on its way. trace: classify ok -> lookup_order failed: order DB timeout -> lookup_order ok (try 2) -> draft_refund ok HANDED TO A HUMAN trace: classify ok -> lookup_order failed: order DB timeout -> lookup_order ok (try 2) -> escalate ok We are open 9 to 5, Monday to Friday. trace: classify ok -> faq ok
Now change it:
- Call
handle(email, retries=0)in the final loop. Predict the reply and trace for each of the three emails. (Keep in mind that the flaky tool counts calls across emails.) - Raise the approval limit from
200to300. Predict which reply changes. Who in a real company should own that number: the prompt, the code, or a config file? - Send the email
"refund order 9999". Predict what happens before you run it. Then decide which kind of failure this is (passing or repeating) and whether more retries would help.
Pause and think: In the second email the lookup succeeded on its second try, yet the request still went to a human. Which line made that choice, and why is it right that code makes it rather than the drafting model?
The elif line: the lookup worked, but the amount (250) is over the limit of 200, so the condition is false and the else branch escalates. A spending limit is a business rule that must hold every time. In code it is checked the same way on every request and can be tested; a model asked to “be careful with large refunds” would follow it most of the time, which is not good enough for money.
Pause and think: The trace for the first email shows a failure and then a success. If we kept only the final reply and threw the trace away, what would we lose?
We would never learn that the order database times out on every other call. The customer got a correct answer, so nothing looks wrong from outside, but each request is slower and the system is one more failure away from escalating. Traces are how we see problems that retries are quietly hiding, and how we find the slow or flaky step in a long workflow.
Key takeaways
- AI orchestration coordinates models, tools, data, memory and agents into one reliable workflow.
- The orchestrator can be code (a workflow) or an LLM (an agent); most real systems mix both.
- Five patterns cover most systems: sequential, parallel, conditional, loop and orchestrator-worker.
- Sequential time adds up; parallel time is the slowest step; loops need hard limits.
- Start simple, make state explicit, trace every step and put humans at the risky points.
Key terms
- AI orchestration: Coordinating LLM calls, tools, data, memory and agents so they complete a task together.
- Orchestrator: The component (code or an LLM) that decides which step runs next and with what inputs.
- Workflow: A system whose steps and branches are fixed in code, with LLMs inside the steps.
- Routing: Choosing which branch, model or agent should handle an input.
- Fan-out / fan-in: Sending work to several parallel steps, then gathering and merging their results.
- Evaluator-optimizer loop: Generate, check, revise, repeated until a check passes or a limit is hit.
← 11.13 Agent Communication: Protocols and Message Formats · 11.15 Sakana Fugu: Lessons from an Open-Source Agent Study →