Lesson 9.2 · 22 min
Prompt Chaining: Decomposing Complex Tasks into Steps
What if, instead of asking one giant prompt to do five jobs at once, we let the model do one job at a time and check its work in between?
In short: Prompt chaining splits a complex task into a sequence of smaller LLM calls, where the output of one call becomes the input of the next. Between calls, ordinary code can validate, route or transform the data. Chains are easier to debug, test and control than one big prompt, at the price of more calls, more latency and the risk that an early mistake flows downstream.
What is a prompt, and what is prompt chaining?
A prompt is the text we send to a large language model (LLM): instructions, examples, data and the question. The model reads the prompt and generates a response, one token (a small piece of text) at a time. One prompt in, one response out: that is a single LLM call.
Prompt chaining is a design pattern where we break a task into several steps and give each step its own LLM call with its own focused prompt. The output of step 1 is placed into the prompt of step 2, and so on. In between, our own code can check, reshape or branch on the results. The whole sequence is called a chain.
Think of it like an assembly line A car factory does not ask one worker to build a whole car. One station fits the frame, the next adds the engine, an inspector checks the welds, and only then does the car move on. Each station has one clear job and a quality check. A prompt chain is an assembly line for text.
Our running example: an online shop receives a customer review: “Ordered the X200 blender on May 2. It arrived cracked and support never replied. I want a refund.” We want the system to understand it, decide what kind of case it is, and draft a polite reply.
Why do we need prompt chaining?
We could write one long prompt: “Read this review, extract the product and issue, decide whether it is a refund, return or question, write a reply under 40 words in our brand voice, and output JSON.” Modern models often manage it. But as the instruction list grows, several problems appear:
- Instructions compete. With many requirements in one prompt, the model tends to satisfy most of them and quietly drop one, and which one it drops can change from run to run.
- Hard to debug. When the final reply is wrong, we cannot tell whether extraction, classification or writing failed.
- No checkpoints. We cannot stop and validate an intermediate result, because there is none: everything happens in one response.
- One size for everything. Every sub-task uses the same model, temperature and prompt, even though extraction wants precision and writing wants fluency.
Chaining fixes these by giving each sub-task a short, specific prompt and by placing plain code between the steps. The model does the language work; code does the checking, which code is very good at.
Pause and think: Pause and predict: a one-prompt pipeline sometimes writes a lovely reply that promises a refund for a customer who only asked a question. With a chain, where would we catch this?
At the classification step. Because classification is its own call with its own output (“refund”, “return” or “question”), code can branch on it and only run the refund-reply prompt when the label is “refund”. We can also log and test that step on its own.
How does prompt chaining work, step by step?
Building and running a chain
- Decompose the task: Write down the sub-tasks a careful human would do in order. For our review: extract facts → classify the case → draft a reply → check the reply.
- Write one focused prompt per step: Each prompt does one thing and asks for an output format the next step can use, often JSON or a single label.
- Pass outputs forward: Insert the previous step's output into the next prompt, usually inside clear delimiters such as
<facts>...</facts>. Pass only what the next step needs. - Add gates between steps: A gate is a code check: is this valid JSON? Are the required fields present? Is the label one of the allowed values? If not, retry the step or stop.
- Branch when needed: Code can choose the next prompt based on a result, for example a refund template vs a question template.
- Log every step: Store each step's input and output so we can find which link broke when something goes wrong.
A real example of prompt chaining
Here are the actual prompts a team might write for the chain above. Notice how small each one is.
| Step | Prompt (shortened) | Output |
|---|---|---|
| 1. Extract | “From the review in <review> tags, return JSON with keys product, issue, request, sentiment. Use null if missing.” | {"product": "X200 blender", "issue": "arrived cracked", ...} |
| 2. Classify | “Given these facts, answer with exactly one word: refund, return or question.” | refund |
| 3. Draft | “Write a reply under 40 words, apologetic and specific, using these facts. Do not promise dates.” | A short reply |
| 4. Check | Code: word count, banned phrases, required facts present. Optionally an LLM check for tone. | pass / fail |
Another everyday chain is summarize a long report: step 1 splits it into sections and summarizes each one, step 2 merges the section summaries, step 3 rewrites the merged summary for a specific audience (for example executives). Each step is easy to inspect on its own.
Where chains show up in real products Document pipelines (extract → validate → store), content tools (outline → draft → edit → fact-check), coding assistants (plan → write code → run tests → fix), translation with review (translate → back-translate → compare), and RAG systems (rewrite the query → retrieve → answer → check citations) are all prompt chains.
Code example of prompt chaining
To keep it runnable anywhere, llm() is a stand-in that returns what a real model would plausibly return. In production we would replace its body with an API call. Everything else, the gate, the branch and the check, is exactly what we would write for real.
review_chain.py
import json
review = "Ordered the X200 blender on May 2. It arrived cracked and support never replied. I want a refund."
def llm(task, text):
# Stand-in for a real model call: each task returns what a model might
if task == "extract":
return json.dumps({"product": "X200 blender", "issue": "arrived cracked",
"request": "refund", "sentiment": "negative"})
if task == "classify":
return "refund" if json.loads(text)["request"] == "refund" else "other"
if task == "draft":
d = json.loads(text)
return (f"Sorry your {d['product']} {d['issue']}. "
f"We have started a {d['request']} and will email you today.")
def gate_valid_json(text, keys):
# A plain-code check between steps: stop early on bad output
try:
data = json.loads(text)
except json.JSONDecodeError:
return False
return all(k in data for k in keys)
step1 = llm("extract", review)
print("Step 1 (extract):", step1)
if not gate_valid_json(step1, ["product", "issue", "request"]):
raise SystemExit("Gate failed: retry step 1")
print("Gate: valid JSON with all keys")
route = llm("classify", step1) # output of step 1 is input of step 2
print("Step 2 (classify):", route)
if route == "refund": # branch on the result
reply = llm("draft", step1)
print("Step 3 (draft):", reply)
n = len(reply.split())
print(f"Step 4 (check): {n} words, under 40? {n <= 40}")Output:
Step 1 (extract): {"product": "X200 blender", "issue": "arrived cracked", "request": "refund", "sentiment": "negative"}
Gate: valid JSON with all keys
Step 2 (classify): refund
Step 3 (draft): Sorry your X200 blender arrived cracked. We have started a refund and will email you today.
Step 4 (check): 16 words, under 40? TrueCommon patterns in prompt chaining
Most real chains are built from a handful of shapes:
- Sequential pipeline. A → B → C. Each step transforms the previous output (extract → classify → draft).
- Gate (validation). Code checks an output and decides: continue, retry, or stop.
- Routing (branching). A classification step picks which prompt or which model runs next.
- Parallel fan-out, then merge. Run the same prompt on many pieces at once (summarize 10 chapters), then one call combines the results. This is often called map-reduce.
- Generate → critique → revise. One call drafts, a second call lists problems, a third call fixes them. Loop a fixed number of times or until a check passes.
Advantages of prompt chaining
- Accuracy. Each call has one goal, so the model gives it its full attention.
- Debuggability. Logs show exactly which step produced a bad output.
- Testability. We can build a small test set for each step (for example 50 reviews with known labels for the classifier).
- Control. Code between steps enforces rules the model might ignore: allowed labels, length limits, required fields.
- Cost tuning. Cheap, fast models can handle easy steps (classification) and a stronger model only the hard ones (writing).
- Reuse. A good extraction step can feed several different downstream chains.
Things to take care of while using prompt chaining
Errors flow downstream If step 1 extracts the wrong product, every later step will confidently build on that mistake. The chart above shows how quickly per-step errors compound. Put a gate after any step whose output later steps depend on, and fail loudly instead of passing bad data along.
- Latency adds up. Sequential calls run one after another. Run independent steps in parallel and use small models for simple steps.
- Cost adds up. Every call re-sends its own instructions and inputs. Keep each prompt lean and pass only what the next step needs.
- Lost context. A later step only knows what we hand it. If the drafting step needs the customer's name, it must be in the hand-off.
- Brittle formats. Ask for structured output (JSON with fixed keys, or one label from a list) and validate it. Many APIs offer a structured-output or JSON mode that helps.
- Too many steps. Splitting a task into ten tiny calls adds latency and hand-offs without improving quality. Split where a human would naturally check the work.
- Test the whole chain too. Each step can pass its own tests while the end-to-end result is still poor. Keep an end-to-end test set.
Pause and think: Our 3-step chain is accurate but too slow. Steps 2 and 3 both read step 1's output but do not depend on each other. What can we change?
Run steps 2 and 3 in parallel (fan-out) since neither needs the other's result, then merge. We can also move a simple step, like classification, to a smaller and faster model.
When to use prompt chaining
| Situation | Use a chain? | Reason |
|---|---|---|
| Task has clear stages (extract, decide, write) | Yes | Each stage gets a focused prompt and a check |
| We need guarantees on format or rules | Yes | Code gates enforce them between steps |
| Different stages need different models | Yes | Route cheap steps to cheap models |
| A single prompt already works reliably | No | Extra calls only add latency and cost |
| Steps cannot be known in advance | Probably not | An agent loop fits better |
| Real-time chat where every 100 ms matters | Carefully | Keep the chain short or parallel |
A good workflow is to start with one prompt, measure where it fails, and split out exactly the failing part into its own step with its own check. That way every link in the chain earns its place.
Worked example, step by step
The chart earlier showed how fast an unchecked chain decays. Let us put numbers on the cure. Take our 4-step review chain and suppose each step is right 90% of the time. The numbers are illustrative, and we assume that tries fail independently.
What a gate with a retry buys us
- No gates: All four steps must be right: 0.9⁴ ≈ 0.66. About one review in three ends with a wrong reply.
- Add a gate with one retry: A step now fails only if both tries fail: 0.1 × 0.1 = 0.01. Its effective success rate is 1 − 0.01 = 0.99.
- Chain the gated steps: 0.99⁴ ≈ 0.96. Same prompts, same model, far fewer bad replies.
- Count the cost: Each step makes 1 call, plus a second call in the 10% of cases where the first try fails: 1.1 calls on average. Four steps cost about 4.4 calls instead of 4.
| Design | Success per step | Whole chain | Average calls |
|---|---|---|---|
| No gates | 0.90 | 0.9⁴ ≈ 0.66 | 4.0 |
| Gate + 1 retry | 0.99 | 0.99⁴ ≈ 0.96 | 4.4 |
| Gate + 2 retries | 0.999 | 0.999⁴ ≈ 0.996 | 4.44 |
Two honest limits. First, a gate catches only what it can check. Invalid JSON is easy to catch. Valid JSON with the wrong product name is not, so real gains are smaller than this table. Second, retries cure random slips. If a step fails the same way every time, a retry repeats the failure, and we must fix the prompt instead.
To tell the two apart, log the gate result for every attempt. Failures that pass on the second try are random slips. Failures that never pass point at the prompt or at the input.
Practice: try it yourself
We will build the report chain from earlier: summarize three sections one by one (fan-out), check each summary with a gate that retries, then merge the summaries in one last call. The model is a scripted stand-in that misbehaves once, so we can watch the gate catch it.
practice_fanout_gate.py
# A chain with fan-out, a retrying gate, a merge step and a call counter.
calls = {"n": 0}
def llm(task, text):
# Scripted stand-in for a model. Its 2nd call ignores the format once.
calls["n"] += 1
if task == "summarize":
if calls["n"] == 2:
return "Sure! Here is a long friendly answer that forgets the format we asked for"
return "SUMMARY: " + text.split(".")[0].lower()
if task == "merge":
return "REPORT: " + "; ".join(s.replace("SUMMARY: ", "") for s in text)
def gate(out):
# Plain-code check: the right prefix and at most 12 words
return out.startswith("SUMMARY: ") and len(out.split()) <= 12
def step_with_retry(task, text, tries=3):
for attempt in range(1, tries + 1):
out = llm(task, text)
ok = gate(out)
print(f" {task} attempt {attempt}: {'pass' if ok else 'FAIL'}")
if ok:
return out
raise SystemExit("gate failed every time: stop the chain")
sections = ["Sales rose in the third quarter. Most growth came from new shops.",
"Support tickets fell. The new help page worked.",
"Two engineers joined. Hiring is on plan."]
summaries = [step_with_retry("summarize", s) for s in sections] # fan-out
report = llm("merge", summaries) # merge
print(report)
print("model calls:", calls["n"])Output:
summarize attempt 1: pass summarize attempt 1: FAIL summarize attempt 2: pass summarize attempt 1: pass REPORT: sales rose in the third quarter; support tickets fell; two engineers joined model calls: 5
Now change it:
- Change
tries=3totries=1. Predict what the program prints and where it stops. - Tighten the gate from 12 words to 4 words. Predict which summaries fail now. What does that say about a gate that is too strict?
- Make the scripted model misbehave on every even-numbered call (
calls["n"] % 2 == 0) instead of only the second. Predict the total number of model calls.
Pause and think: The output shows 5 model calls for 3 sections plus 1 merge. Where did the fifth call come from, and what would have happened without the gate?
From the retry. The second summarize call ignored the format, the gate rejected it, and the step ran again. Without the gate, that chatty text would have gone into the merge step and ended up inside the final report. The gate spent one extra call to stop an error at its source.
Pause and think: The gate checks the prefix and the word count. Suppose a summary says “sales fell” when the section says sales rose. Does the gate catch it? What kind of check could?
No. The gate checks form, not truth. The wrong summary has the right prefix and length, so it passes. Catching it needs a check on content: a second model call that compares the summary with its section, or a test set with known answers for this step. Format gates are cheap and worth having, but they do not replace measuring each step's accuracy.
Key takeaways
- Prompt chaining splits a task into focused LLM calls; each output becomes the next input.
- Code between steps (gates) validates, routes and transforms, catching errors before they spread.
- Common shapes: sequential pipeline, gate, routing, parallel fan-out/merge, and generate-critique-revise.
- Chains are easier to debug and test than one big prompt, but cost more calls and latency.
- Use a chain when the steps are known in advance; use an agent when they are not.
Key terms
- LLM call: One request to a language model: a prompt in, one response out.
- Prompt chaining: Running several focused LLM calls in sequence, passing each output into the next prompt.
- Gate: A code check between chain steps that decides to continue, retry or stop.
- Routing: Choosing the next prompt or model based on the result of a classification step.
- Fan-out / merge: Running a step on many pieces in parallel, then combining the results in one call.
- Error propagation: A mistake in an early step being carried into and amplified by later steps.
← 9.1 Chain-of-Thought Prompting: Making Models Reason Step by Step · 9.3 Prompt Caching: Reusing Computation Across API Calls →