Lesson 9.1 · 23 min
Chain-of-Thought Prompting: Making Models Reason Step by Step
Why does adding one sentence, “Let's think step by step”, make a language model noticeably better at word problems?
In short: Chain-of-Thought (CoT) prompting asks a language model to write out its intermediate reasoning before giving the final answer. Because every generated token is extra computation the model can build on, writing the steps down turns one hard leap into many small, easier steps. We can trigger it with a short instruction (zero-shot) or with worked examples (few-shot), and we can make it more reliable by sampling several chains and taking a vote.
First, what is a prompt and what is an LLM?
A Large Language Model (LLM) is a neural network trained on a huge amount of text to do one job: given some text, predict the next token. A token is a small piece of text, often a word or part of a word. To write a full answer, the model predicts one token, appends it to the text, predicts the next one, and repeats. This is called autoregressive generation.
A prompt is the text we give the model as its starting point: our question, instructions, examples and any data it needs. The model never sees anything except the prompt plus the tokens it has already written. So the way we phrase the prompt shapes what the model writes next. Prompt engineering is the craft of writing prompts that reliably get good answers.
Think of it like a student with scratch paper Ask a student “What is 17 × 24?” and demand an answer in one second, and they may guess. Give them scratch paper and they write 17 × 20 = 340, 17 × 4 = 68, 340 + 68 = 408. Same student, same knowledge, better answer. Chain-of-Thought prompting is how we hand the model scratch paper.
Our running example in this lesson is a small word problem: “A cafe has 23 apples. It uses 20 for lunch and buys 6 more. How many apples does it have now?” The right answer is 23 − 20 + 6 = 9.
The problem: when the model jumps straight to the answer
If our prompt says “Answer with just a number”, the very first token the model writes must already be the answer. That means the model has to do all the arithmetic inside one forward pass of the network, with no room to write anything down.
Here is the key fact: a Transformer spends roughly the same amount of computation on every token it generates. A fixed number of layers runs once per token. An easy question and a hard question both get the same budget for that single answer token. For a multi-step problem, that budget can be too small, so the model falls back on a pattern that looks right, such as copying a number from the question or doing only one of the two operations.
- It might answer 29 (23 + 6, forgetting the 20 used for lunch).
- It might answer 3 (23 − 20, forgetting the 6 bought).
- It might answer 9, but we cannot see why, so we cannot check it.
This failure mode is common on math word problems, multi-hop questions (“Who was president when the author of X was born?”), logic puzzles and anything where a later step depends on an earlier result.
What is Chain-of-Thought prompting?
Chain-of-Thought (CoT) prompting is a prompting technique in which we get the model to produce a sequence of intermediate reasoning steps (the “chain of thought”) before the final answer. The technique was described and named by Wei and colleagues at Google in the 2022 paper Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.
Nothing about the model changes. No training, no new weights. We only change the prompt so that the most likely continuation is a worked solution rather than a bare number.
Two things improved. The answer is more likely to be right, because each step is small. And the answer is inspectable: if it were wrong, we could see exactly which line went wrong.
Pause and think: Pause and predict: if CoT does not change the model's weights, where does the extra accuracy come from?
From extra computation at inference time. Every reasoning token is another full forward pass whose result is written into the context, and later tokens can read it. The model gets to store intermediate results (like “3”) instead of holding everything in one pass.
Zero-shot CoT vs Few-shot CoT
There are two classic ways to trigger a chain of thought. Few-shot CoT (the original Wei et al. method) puts a few worked examples, each with question, reasoning and answer, in the prompt. The model copies that format for the new question. Zero-shot CoT (Kojima et al., 2022, Large Language Models are Zero-Shot Reasoners) adds no examples at all, just a trigger phrase like “Let's think step by step.”
A practical variant asks for the reasoning and the answer in separate, labelled parts, for example “Think in a section called Reasoning, then write Answer: <number>.” This makes it easy for code to extract the answer.
A step-by-step walkthrough of a reasoning chain
Let us follow what happens, token by token, when we send the CoT prompt for our cafe problem.
One chain of thought, from prompt to answer
- Read the prompt: The model processes the question plus “Let's think step by step.” The likeliest continuation is now a worked explanation, not a number.
- Restate the start: It writes “The cafe starts with 23.” This copies the key quantity into a fresh, nearby spot in the context.
- First operation: It writes “23 − 20 = 3”. This is one small subtraction, easy to get right in a single step. The result “3” is now in the context.
- Second operation: It writes “3 + 6 = 9”. It reads the “3” it wrote a moment ago instead of recomputing it.
- State the answer: It writes “The answer is 9.” Our code finds this line with a regular expression and grades it.
Why does Chain-of-Thought prompting work?
There is no single proven explanation, but several reasons are widely accepted:
- More computation per answer. Each token gets a fixed amount of compute. Writing 40 reasoning tokens before the answer means 40 more forward passes are spent on the problem.
- A working memory. Intermediate results are written into the context, so later steps can attend to them. The model does not have to carry everything in its hidden state.
- Decomposition. A hard problem becomes a chain of easy sub-problems, each of which looks like something the model has seen many times in training.
- It matches the training data. The web is full of worked solutions, tutorials and step-by-step explanations. The trigger phrase pushes the model into that familiar style.
The original paper also observed that the benefit depends on model size: in their experiments CoT helped large models a lot, while for small models it gave little benefit or even hurt, because small models wrote fluent but wrong reasoning. Today's strong models handle CoT well.
Code you can run: building CoT prompts and voting
We cannot call a real model here, so we hard-code five chains that a model might sample. The code shows the parts we really write in production: the two prompt styles, an answer extractor, and a self-consistency vote.
cot_vote.py
import re
from collections import Counter
question = "A cafe has 23 apples. It uses 20 for lunch and buys 6 more. How many apples now?"
# Zero-shot CoT: append a trigger phrase, no examples
zero_shot = f"Q: {question}\nA: Let's think step by step."
# Few-shot CoT: show one worked example (with its reasoning) first
demo = ("Q: Tom has 5 pens and buys 2 packs of 3 pens. How many pens?\n"
"A: He starts with 5. Two packs of 3 is 6. 5 + 6 = 11. The answer is 11.")
few_shot = f"{demo}\n\nQ: {question}\nA:"
# Pretend we sampled 5 reasoning chains from a model at temperature 0.7
chains = [
"Start with 23. 23 - 20 = 3. 3 + 6 = 9. The answer is 9.",
"23 apples minus 20 is 3. Buying 6 gives 9. The answer is 9.",
"23 - 20 = 13. 13 + 6 = 19. The answer is 19.", # an arithmetic slip
"After lunch: 3 left. Plus 6 bought: 9. The answer is 9.",
"23 + 6 = 29. 29 - 20 = 9. The answer is 9.",
]
def final_answer(chain):
# Pull out the number after 'The answer is' so code can grade it
m = re.search(r"The answer is (-?\d+)", chain)
return int(m.group(1)) if m else None
answers = [final_answer(c) for c in chains]
best, count = Counter(answers).most_common(1)[0]
print("Zero-shot CoT prompt:\n" + zero_shot)
print("\nFew-shot CoT prompt:", len(few_shot.split()), "words")
print("Extracted answers:", answers)
print(f"Self-consistency vote: {best} ({count} of {len(chains)} chains)")Output:
Zero-shot CoT prompt: Q: A cafe has 23 apples. It uses 20 for lunch and buys 6 more. How many apples now? A: Let's think step by step. Few-shot CoT prompt: 55 words Extracted answers: [9, 9, 19, 9, 9] Self-consistency vote: 9 (4 of 5 chains)
Pause and think: In the output, one chain answered 19. Why did the final answer still come out as 9?
Because we voted. Four chains independently reached 9 and only one reached 19. Correct reasoning paths tend to converge on the same answer, while mistakes scatter, so the majority answer is more reliable than any single chain.
Where Chain-of-Thought prompting is useful
| Task | Does CoT help? | Why |
|---|---|---|
| Math word problems | Usually a lot | Several dependent arithmetic steps |
| Multi-hop questions | Yes | Each hop's answer feeds the next hop |
| Logic and planning puzzles | Yes | Constraints must be checked one by one |
| Code debugging | Often | Tracing values line by line finds the bug |
| Agents choosing tools | Yes | Patterns like ReAct interleave reasoning with tool calls |
| Simple fact lookup (“capital of France?”) | Little or none | One step; extra tokens only add cost |
| Sentiment of a short review | Little | Pattern recognition, not multi-step reasoning |
Real-world use A tutoring app asks the model to solve a problem step by step, then checks the final answer against a known key, and only shows the student the steps if the answer matched. A support bot reasons privately about which refund rule applies, then shows the customer only the short conclusion. Agent frameworks use reasoning steps before each tool call so the next action is planned, not guessed.
Things to keep in mind
The reasoning is not a window into the model's mind A chain of thought is generated text, not a log of the network's internal computation. Research has shown that models can produce plausible reasoning that does not reflect what actually drove the answer (the “faithfulness” problem). Treat the chain as a useful aid and a debugging hint, not as proof.
- Cost and latency. Reasoning tokens are output tokens, which are billed and take time. A 5-token answer can become 200 tokens.
- Errors can cascade. A wrong early step is copied into every later step. Self-consistency or a verification step helps.
- Parse the answer. Ask for a fixed final line (
Answer: ...) so code does not have to guess which number in the prose is the answer. - Do not show raw reasoning to end users by default. It can be long, confusing, or contain wrong intermediate claims. Show the conclusion.
- Reasoning models change the picture. Models trained to think before answering (for example OpenAI's o-series, DeepSeek-R1, and the extended-thinking modes of Claude and Gemini) already produce a hidden or visible chain of thought. Adding “think step by step” to them usually adds little, and some vendors advise against prescribing the steps. Give them a clear goal instead.
- Skip it for one-step tasks. If the task has no intermediate steps, CoT costs more and may even overthink.
Common mistakes and how to spot them
Most CoT failures look the same from the outside: a wrong final answer. The written chain is what lets us tell them apart. Here are three patterns we meet often, shown on our cafe problem.
| Mistake | What we see | Why it happens | Fix |
|---|---|---|---|
| Answer first, reasons after | A number, then a tidy explanation that may not match it | The answer was written before any step existed, so the steps could not help it | Ask for the steps first and the answer on the last line |
| Cascading slip | 23 − 20 = 13, then 13 + 6 = 19 | Every later step trusts the earlier line | Re-check each line with code, or vote over several chains |
| Wrong number parsed | The chain is right, but our code returns 23 or 6 | The parser grabs the first number it finds, not the answer line | Demand a fixed last line such as Answer: 9 and parse only that |
The first row surprises many people. A model writes from left to right. Text that comes after the answer cannot change the answer. So a prompt that says “give the answer, then explain” gets almost none of the benefit of CoT, even though the output looks like reasoning.
Diagnosing one wrong chain
- Find the answer line: Check that our code parsed the line we meant. If it did not, we have a parsing bug, not a reasoning bug.
- Re-compute each step: Walk down the chain and redo each small operation. The first line that does not hold is where the error began.
- Pick the fix: A slip that changes from run to run calls for a vote or a verification step. The same misreading on every run calls for a clearer prompt or a worked example.
Practice: try it yourself
We will build a toy “model” that can do only one arithmetic operation per pass. Asked for a direct answer, it runs out of room after the first operation. Allowed to write each result on a scratchpad and read it back, it reaches the multi-step answer. We also add a checker that finds the first wrong line in a chain.
practice_scratchpad.py
# Toy model: each "forward pass" can do only ONE arithmetic operation.
OPS = {"+": lambda a, b: a + b, "-": lambda a, b: a - b, "*": lambda a, b: a * b}
def one_pass(start, steps):
# Direct answer: a single pass, so only the first operation gets done
op, n = steps[0]
return OPS[op](start, n)
def chain_of_thought(start, steps, slip_at=None):
# One operation per pass; each result is written down and read back
pad, value = [], start
for i, (op, n) in enumerate(steps):
result = OPS[op](value, n)
if i == slip_at:
result += 10 # inject an arithmetic slip
pad.append((value, op, n, result)) # one scratchpad line
value = result # the next pass reads this line
return value, pad
def first_bad_line(pad):
# The steps are written down, so code can re-check each one
for i, (a, op, n, result) in enumerate(pad, 1):
if OPS[op](a, n) != result:
return i
return None
steps = [("-", 20), ("+", 6), ("*", 2)] # 23 apples: use 20, buy 6, double
print("direct answer:", one_pass(23, steps), "(the truth is 18)")
answer, pad = chain_of_thought(23, steps)
for a, op, n, r in pad:
print(f" {a} {op} {n} = {r}")
print("chain answer:", answer, "| passes used:", len(pad))
answer, pad = chain_of_thought(23, steps, slip_at=1)
print("with a slip:", answer, "| first bad line:", first_bad_line(pad))Output:
direct answer: 3 (the truth is 18) 23 - 20 = 3 3 + 6 = 9 9 * 2 = 18 chain answer: 18 | passes used: 3 with a slip: 38 | first bad line: 2
Now change it:
- Add a fourth step,
("-", 4), tosteps. Before running, predict the chain answer and the number of passes. Does the direct answer change? - Change
slip_at=1toslip_at=0. Predict the final answer and the line thatfirst_bad_linereports. - Give
one_passabudgetargument so it applies the firstbudgetoperations. Predict the smallest budget that returns 18, and say what that budget stands for in a real model.
Pause and think: In the slip run, the second and third scratchpad lines both hold wrong values (19 and 38), yet first_bad_line reports only line 2. Why is line 3 not flagged?
Line 3 is 19 * 2 = 38, which is correct arithmetic on a wrong input. The checker tests each step on its own, so it flags only the step where the error was made. This is what a cascade looks like: one bad step, then correct steps that carry the bad value forward. Fix the first bad line and the later lines repair themselves.
Pause and think: A teammate changes the prompt to “Reply with the final number first, then show your working.” The working looks neat, but accuracy falls back to the no-CoT level. What happened?
The answer is now generated before any step is written. A model writes left to right, so the working cannot feed into a number that is already on the page. The extra computation helps only when the steps come first and the answer comes last.
Key takeaways
- CoT prompting makes the model write intermediate steps before the answer; only the prompt changes.
- Each generated token is extra computation and stored working memory, so many small steps beat one big leap.
- Zero-shot CoT uses a trigger phrase; few-shot CoT uses worked examples for more control.
- Self-consistency samples several chains and takes a majority vote to cancel out random slips.
- CoT costs more tokens, can be unfaithful, and adds little for one-step tasks or built-in reasoning models.
Key terms
- Prompt: The text we give an LLM as its starting point: instructions, examples, data and the question.
- Chain-of-Thought (CoT): A prompting technique where the model writes intermediate reasoning steps before its final answer.
- Zero-shot CoT: Triggering reasoning with an instruction such as “Let's think step by step”, without examples.
- Few-shot CoT: Triggering reasoning by including worked examples that show the reasoning format.
- Self-consistency: Sampling several reasoning chains and returning the most common final answer.
- Faithfulness: Whether a model's written reasoning truly reflects the process that produced its answer.
← 8.11 GRPO: Group-Based Preference Optimization Explained · 9.2 Prompt Chaining: Decomposing Complex Tasks into Steps →