Modern AI Engineering

Lesson 11.5 · 23 min

Plan-and-Execute: Tackling Complex Tasks in Two Phases

Should an agent decide each move as it goes, or think through the whole route before taking the first step?

In short: A Plan-and-Execute agent splits the work in two: a planner (usually a strong LLM) writes the full list of steps up front, and an executor carries them out one by one, often with a cheaper model or plain code. When a step fails or reveals something new, a replanner revises the remaining plan. This gives structure, lower cost and easier oversight on long tasks, at the price of less flexibility than step-by-step agents like ReAct.

What is a Plan-and-Execute agent?

A Plan-and-Execute agent solves a task in two separate phases:

  • Plan: one model call reads the goal and writes an ordered list of steps (the plan), before anything is executed.
  • Execute: each step is carried out in order, using tools. Results are collected as the steps complete.

A third, optional component makes it robust: a replanner that looks at what has happened so far and rewrites the remaining steps when reality does not match the plan (a step failed, or a result changes what makes sense next).

The idea draws on research such as Plan-and-Solve prompting (Wang and colleagues, 2023), which showed that asking a model to first devise a plan and then carry it out can improve multi-step reasoning, and on open-source agent projects of 2023 that kept an explicit task list. Agent frameworks such as LangChain and LangGraph later offered it as a standard pattern.

Think of it like a building project An architect (the planner) draws the full blueprint before construction starts. Builders (the executors) follow it step by step without redesigning the house at each brick. When the builders hit rock where the foundation should go, they call the architect back (the replanner), who revises the remaining drawings. Compare that with a builder who decides each brick as they go: flexible, but slow, and the house may not end up coherent.

Plan-and-Execute agent vs a basic AI agent

A basic tool-calling agent has no explicit plan. At each turn the model looks at the history and picks one next action. Any planning lives implicitly in the model's head and is re-done every turn. A Plan-and-Execute agent makes planning a separate, visible artifact: a list we can read, log, show to a user, edit, or approve before execution starts.

What changes when the plan becomes explicit
AspectBasic agent loopPlan-and-Execute
When planning happensImplicitly, every turnExplicitly, once up front (plus replans)
Can a human review the plan before acting?Not easilyYes, the plan is a list
Model used for each stepThe same model every turnStrong planner; cheaper model or code for steps
Sense of overall progressHidden in the historyClear: step 3 of 5 done

Pause and think: Why is it easier to put a human approval gate in a Plan-and-Execute agent than in a basic loop?

Because the whole plan exists as a readable list before any tool runs. A person can approve, edit or reject it once, up front. In a basic loop, actions are decided one at a time, so approval would have to happen at every turn.

Anatomy of a Plan-and-Execute agent

The four parts
PartJobTypical choice
PlannerTurn the goal into an ordered list of concrete steps.The strongest (most expensive) model, called once.
ExecutorCarry out one step at a time with tools; return the result.A smaller model running a short tool loop, or plain code.
StateHold the goal, the remaining plan, and results of completed steps.A simple object or dictionary passed between parts.
ReplannerAfter a step (or on failure), keep, edit, or extend the remaining plan, or finish.The planner model again, given the results so far.

Make the plan structured Ask the planner for a JSON list (use structured outputs if your API supports them), where each step has an id, a description, the tool it expects to use, and which earlier steps it depends on. Structured plans are easy to validate, display and execute.

How a Plan-and-Execute agent works

From goal to answer

  1. Receive the goal: For example: “Plan a one-night trip from London to Nice, as cheap as possible.”
  2. Plan: The planner writes the steps: 1) find a flight London→Nice, 2) find a hotel in Nice. Optionally a human approves.
  3. Execute the next step: The executor runs step 1 with the flight tool and stores the result.
  4. Check and replan: If the step succeeded, move on. If it failed (no direct flight), the replanner rewrites the remaining steps: fly to Paris, then take a train to Nice.
  5. Repeat until done: Execute the remaining steps in order, replanning only when needed.
  6. Synthesise: With all results in hand, write the final answer: the route, the hotel, and the total cost.

Because steps are known in advance, independent steps can run in parallel. If the plan says “search flights” and “search hotels” and neither depends on the other, the executor can do both at once. Research systems such as LLMCompiler took this further by planning a dependency graph of tool calls and running independent ones concurrently.

A full trace example

Let us run the trip example. The planner's first plan assumes a direct flight. In our toy data there is none, so step 1 fails and the replanner routes through Paris. Notice that the hotel step from the original plan is kept: the replanner only replaces what broke.

plan_and_execute.py

# Plan-and-execute: planner writes all steps up front; executor runs them;
# a replanner patches the plan when a step fails.
FLIGHTS = {"LHR->CDG": 120}             # toy data: no direct LHR->NCE flight here
TOOLS = {
"flight": lambda r: FLIGHTS.get(r) or f"ERROR no flight {r}",
"train":  lambda r: {"CDG->NCE": 80}.get(r, f"ERROR no train {r}"),
"hotel":  lambda city: {"NCE": 95}[city],
}
def planner(goal):                       # one "big" LLM call -> a list of steps
return [("flight", "LHR->NCE"), ("hotel", "NCE")]
def replanner(failed_step, error):       # called only when something breaks
print(f"  replan: {failed_step} failed ({error}); route via Paris")
return [("flight", "LHR->CDG"), ("train", "CDG->NCE")]
plan = planner("Trip London -> Nice, cheapest, 1 night")
print("initial plan:", plan)
done, costs, replans = [], [], 0
while plan:
step = plan.pop(0)                   # executor: cheap model or plain code
result = TOOLS[step[0]](step[1])
if isinstance(result, str) and result.startswith("ERROR") and replans < 2:
replans += 1
plan = replanner(step, result) + plan   # keep the remaining steps
continue
done.append(step); costs.append(result)
print(f"  ran {step} -> {result}")
print("total cost:", sum(costs), "GBP | steps run:", len(done), "| replans:", replans)

Output:

initial plan: [('flight', 'LHR->NCE'), ('hotel', 'NCE')]
replan: ('flight', 'LHR->NCE') failed (ERROR no flight LHR->NCE); route via Paris
ran ('flight', 'LHR->CDG') -> 120
ran ('train', 'CDG->NCE') -> 80
ran ('hotel', 'NCE') -> 95
total cost: 295 GBP | steps run: 3 | replans: 1

Pause and think: In the trace, why was the planner called only once at the start, even though four steps were attempted?

Because execution follows the plan without asking the planner each time. The big model is called once to plan and once more (as the replanner) only when step 1 failed. Steps themselves run with tools or a cheaper executor. This is where the cost savings come from.

Plan-and-Execute agent vs ReAct agent

Common failure modes and how to fix them

Plan-and-Execute failures and fixes
FailureWhat happensFix
Bad initial planSteps are vague (“research the topic”), missing, or in the wrong order.Ask for concrete, tool-sized steps; give example plans; validate the plan's structure.
Stale planA result changes the situation but the executor keeps following the old plan.Run the replanner after each step, or at least after any unexpected result.
Replanning loopsThe replanner keeps rewriting the plan and never finishes.Cap replans; require progress; fall back to a human.
Lost context in stepsThe executor does not know why a step matters or what earlier steps found.Pass the goal and relevant prior results into each executor call.
Over-planningA 2-step task gets a 12-step plan; cost and latency balloon.Ask for the minimum number of steps; skip planning for simple requests.
Unverified completionEvery step “succeeds” but the final result does not meet the goal.A final check step against the original goal before answering.

The most common mistake Treating the plan as fixed. Real environments are full of surprises (a missing file, an empty search result, an API error). Without a replanner, a Plan-and-Execute agent marches confidently through a plan that no longer makes sense.

Real-world use Deep-research features commonly start by drafting a research plan (sometimes shown to the user for approval) and then run many searches against it. Coding agents often write a task checklist before editing files and tick items off as they go. Data pipelines use a plan of queries whose independent parts run in parallel.

Quick summary. A Plan-and-Execute agent separates thinking from doing: a planner writes the steps, an executor carries them out, and a replanner fixes the plan when reality disagrees. It saves strong-model calls, enables parallelism and human review, and keeps long tasks on track. It is less nimble than ReAct, so pair it with replanning and use ReAct-style execution inside steps when needed.

Worked example, step by step

The dependency matrix told us which steps can run together. Let us now work out when each step runs and how much time that saves. We use the same five steps and pretend every step takes 2 seconds (an illustrative number, chosen to keep the sums easy).

The rule is simple: a step may start as soon as every step it depends on has finished. Applying that rule again and again sorts the plan into waves. All steps in one wave run at the same time.

Sorting the plan into waves

  1. Wave 1: steps with no dependencies: Steps 1 (flights), 2 (trains) and 3 (hotels) need nothing. They all start at 0 s and finish at 2 s.
  2. Wave 2: what is unlocked now?: Step 4 (compare routes) needs 1 and 2, which are done. Step 5 needs 3 and 4, and 4 is not done yet. So wave 2 is only step 4, from 2 s to 4 s.
  3. Wave 3: the last step: Step 5 (book) now has both of its inputs. It runs from 4 s to 6 s.
  4. Nothing left: Every step has a result, so the executor stops and hands the results to the final answer.
Run one by one vs run in waves (illustrative: 2 s per step)
Way of runningOrderTotal time
One by one1, 2, 3, 4, 55 × 2 s = 10 s
In waves{1, 2, 3} then {4} then {5}3 × 2 s = 6 s

Notice what sets the total: not the number of steps, but the longest chain of steps that must wait for each other. Here that chain is 1 → 4 → 5 (or 2 → 4 → 5), three steps long, so three waves is the best we can do. Adding ten more independent searches to wave 1 would not add any time; adding one more step after step 5 would.

Two cautions. If two steps wait on each other (4 needs 5 and 5 needs 4), no step ever becomes ready and the executor must detect that and stop. And running steps together is safe for lookups; steps that change things, such as the booking, should stay in their own wave and usually behind an approval.

Practice: try it yourself

The earlier example showed replanning. Here we build the other half: an executor that reads the dependencies in a plan, runs every ready step in a wave, and passes earlier results into later steps. The planner is scripted and is called exactly once.

practice_plan_waves.py

# Run a plan as a dependency graph: steps whose inputs are ready run in one wave.
def planner(goal):
"""Scripted planner: ONE call returns every step, its tool and what it needs."""
return {
"s1": {"tool": "flight",   "arg": "LHR->NCE", "needs": []},
"s2": {"tool": "train",    "arg": "LHR->NCE", "needs": []},
"s3": {"tool": "hotel",    "arg": "NCE",      "needs": []},
"s4": {"tool": "cheapest", "arg": None,       "needs": ["s1", "s2"]},
"s5": {"tool": "total",    "arg": None,       "needs": ["s3", "s4"]},
}
# Each tool gets its own argument plus the results of the steps it depends on.
TOOLS = {
"flight":   lambda arg, inputs: 140,          # made-up prices
"train":    lambda arg, inputs: 110,
"hotel":    lambda arg, inputs: 95,
"cheapest": lambda arg, inputs: min(inputs),
"total":    lambda arg, inputs: sum(inputs),
}
plan = planner("One night in Nice from London, as cheap as possible")
results, wave = {}, 0
while len(results) < len(plan):
# A step is ready when it has not run yet and all its needs have results.
ready = [s for s, step in plan.items()
if s not in results and all(n in results for n in step["needs"])]
if not ready:
print("stuck: the remaining steps wait on each other")
break
wave += 1
outputs = {}
for s in ready:                               # these could run concurrently
step = plan[s]
inputs = [results[n] for n in step["needs"]]   # fill in earlier results
outputs[s] = TOOLS[step["tool"]](step["arg"], inputs)
results.update(outputs)
print(f"wave {wave}: {outputs}")
print(f"planner calls: 1 | steps: {len(results)} | waves: {wave} | answer: {results.get('s5')}")

Output:

wave 1: {'s1': 140, 's2': 110, 's3': 95}
wave 2: {'s4': 110}
wave 3: {'s5': 205}
planner calls: 1 | steps: 5 | waves: 3 | answer: 205

Now change it:

  • Make the hotel depend on the route: set s3 to "needs": ["s4"]. Predict the new waves and the number of waves before running. Does the answer change? (Careful: hotel ignores its inputs.)
  • Create a cycle: set s4 to "needs": ["s1", "s2", "s5"]. Predict exactly which steps run and which line is printed after them.
  • Change the train price to 150. Predict which wave outputs change and what the final answer becomes.

Pause and think: The output says there were 5 steps but only 1 planner call. In a step-by-step agent, roughly how many strong-model calls would the same task take, and what did we give up to save them?

A step-by-step agent would call the model about once per step plus once for the answer, so about 6 calls. We saved them by fixing all five steps up front. What we gave up is the chance to react between steps: if the train search had failed, this executor would have no way to change course. That is the job of the replanner from the earlier example.

Pause and think: Why does the code store a wave's results in outputs and only then copy them into results, instead of writing to results straight away?

Steps in one wave are meant to be independent and could run at the same time. If each step wrote to results immediately, a later step in the same loop could see an earlier one's result, and the behaviour would depend on the order of the loop. Collecting first and updating after the wave keeps every step's inputs limited to finished waves, which is what makes true parallel execution safe.

Key takeaways

  • Plan-and-Execute splits work into a planner (writes all steps), an executor (does them), and a replanner (fixes the plan).
  • Explicit plans are reviewable, show progress, and enable parallel execution of independent steps.
  • The strong model is used mainly for planning, which can cut cost versus calling it at every step.
  • Its weakness is rigidity: without replanning, a plan goes stale when results surprise it.
  • Choose ReAct for exploratory tasks and Plan-and-Execute for structured, decomposable ones; hybrids are common.

Key terms

  • Plan-and-Execute agent: An agent that writes a complete plan first and then executes its steps, revising the plan when needed.
  • Planner: The model call that turns a goal into an ordered list of concrete steps.
  • Executor: The component (smaller model or code) that carries out one plan step with tools.
  • Replanner: The component that reviews progress and rewrites the remaining steps, or declares the task done.
  • Dependency: A relationship where one step needs another step's result before it can run.
  • Stale plan: A plan that no longer fits the situation because results differed from what the planner assumed.

← 11.4 ReAct Agents: Interleaving Reasoning and Acting · 11.6 Reflection Agents: Self-Critique for Higher-Quality Outputs →