Modern AI Engineering

Lesson 7.2 · 27 min

Large Reasoning Models: Chain-of-Thought at Inference Time

Why would a model that answers more slowly, and costs more per question, be the right choice for your hardest problems?

In short: A Large Reasoning Model (LRM) is an LLM trained, mostly with reinforcement learning on problems with checkable answers, to produce a long chain of reasoning before its final answer. Spending more tokens on thinking at answer time (test-time compute) makes it markedly better at maths, coding, logic and planning. The cost is latency and tokens, so we use LRMs for hard, multi-step problems and regular LLMs for simple, fast tasks.

The big picture

A regular LLM answers by predicting the next token, immediately. For "What is the capital of Japan?" that is perfect. For "Find the bug in this 300-line concurrency code" or a tricky maths proof, answering immediately is like blurting out the first thought. In 2024–2025 a new kind of model appeared that first thinks in a long internal monologue, checks itself, backtracks, and only then answers.

OpenAI's o1 (September 2024) was the first widely used example. DeepSeek-R1 (January 2025) showed openly how to train one. Since then most labs offer reasoning models or "thinking" modes. We call them Large Reasoning Models (LRMs).

Think of it like an exam with scratch paper A regular LLM must write the final answer straight onto the answer sheet. An LRM gets scratch paper: it can try an approach, notice an error, try another, and only then copy a clean answer over. On easy questions the scratch paper is a waste of time; on hard ones it makes all the difference.

Running example: a developer tools company with two jobs. Job A: label incoming GitHub issues as bug, feature or question. Job B: given a failing test and the code, find the root cause and propose a fix. We will see why these call for different kinds of models.

What is a Large Reasoning Model (LRM)?

An LRM is still a Transformer language model, usually starting from a strong pre-trained LLM. What changes is training and behaviour: it is trained to generate a long stream of reasoning tokens (also called a chain of thought or "thinking") before the final answer. The reasoning may be shown to the user, summarised, or hidden, depending on the product.

In its thinking, an LRM typically breaks the problem into parts, tries a solution, verifies intermediate results, notices mistakes ("wait, that contradicts step 2"), and explores alternatives. These behaviours were not hand-coded; they emerged or were reinforced because they led to correct answers during training.

LLM vs LRM

How does an LRM actually think?

Mechanically, thinking is just more next-token prediction. The difference is that the model writes intermediate steps into its own context, and every later token can attend to them. Each written step becomes working memory. A Transformer has a fixed amount of computation per token, so producing more tokens is literally giving it more computation for the problem.

The trace is not a perfect window into the model Research has found that reasoning traces do not always faithfully reflect what drove the final answer; a model can reach an answer for reasons it does not write down. Treat the trace as a useful debugging aid, not as proof that the answer is right.

Test-time compute: thinking longer makes them smarter

Test-time compute (or inference-time compute) means spending more computation when answering, rather than only when training. OpenAI reported with o1 that accuracy improved smoothly as the model was allowed to think longer, a new scaling axis alongside model size and training data. There are two main ways to spend it:

  • Sequential: one longer chain of thought, with more steps, checks and revisions. This is what an LRM's "reasoning effort" setting usually controls.
  • Parallel: sample several independent solutions and combine them, for example by majority vote (also called self-consistency) or by a verifier that picks the best one.

more_samples_more_accuracy.py

import numpy as np
rng = np.random.default_rng(0)
# A model solves a problem correctly 60% of the time per attempt.
# Wrong attempts scatter across 4 different wrong answers.
p_correct, n_wrong, trials = 0.6, 4, 20_000
def one_attempt(size):
correct = rng.random(size) < p_correct
wrong_choice = rng.integers(1, n_wrong + 1, size)   # answers 1..4 are wrong
return np.where(correct, 0, wrong_choice)            # answer 0 is right
for k in [1, 3, 5, 9, 17]:
answers = one_attempt((trials, k))                   # k samples per question
votes = np.apply_along_axis(np.bincount, 1, answers, minlength=n_wrong + 1)
votes = votes + rng.random(votes.shape) * 0.1         # break ties at random
majority_right = (votes.argmax(1) == 0).mean()
print(f"samples={k:>2}  majority-vote accuracy={majority_right:.3f}  "
f"tokens spent ~ {k * 800:>6} (800 per attempt)")
# Reward used in RL with verifiable answers: 1 if the final answer matches, else 0
def reward(model_answer, reference):
return 1.0 if model_answer.strip() == reference.strip() else 0.0
print("reward('42', '42') =", reward("42", "42"), "| reward('41', '42') =", reward("41", "42"))

Output:

samples= 1  majority-vote accuracy=0.597  tokens spent ~    800 (800 per attempt)
samples= 3  majority-vote accuracy=0.725  tokens spent ~   2400 (800 per attempt)
samples= 5  majority-vote accuracy=0.836  tokens spent ~   4000 (800 per attempt)
samples= 9  majority-vote accuracy=0.940  tokens spent ~   7200 (800 per attempt)
samples=17  majority-vote accuracy=0.993  tokens spent ~  13600 (800 per attempt)
reward('42', '42') = 1.0 | reward('41', '42') = 0.0

Pause and think: In the simulation, why does voting help so much? When would it fail?

Correct answers agree with each other while wrong answers scatter, so the right answer usually wins the vote. If the model makes the same wrong answer most of the time (a systematic error), voting amplifies that error instead, and more samples do not help.

How are LRMs trained?

The key ingredient is reinforcement learning with verifiable rewards (RLVR). Instead of humans rating answers, the training uses problems whose answers can be checked automatically: maths with known results, code with unit tests, puzzles with a checker. The model generates reasoning and an answer; a program checks the answer and gives a reward; the model is updated to make rewarded reasoning more likely.

A typical LRM training pipeline (details vary by lab)

  1. Start from a strong base LLM: Pre-training gives knowledge and language skill.
  2. Cold-start SFT (optional): Fine-tune on a small set of high-quality long reasoning examples so the model learns the format and readable style.
  3. Large-scale RL on verifiable tasks: Sample several solutions per problem, reward correct final answers (and often a correct format), and update with a policy-gradient method. DeepSeek-R1 used GRPO, which compares each sample's reward with the average of its group instead of training a separate value model.
  4. Broaden and align: More SFT and RL on general tasks so the model stays helpful and safe outside maths and code.
  5. Distil (optional): Use the LRM's reasoning traces to fine-tune smaller models, which become surprisingly good reasoners. DeepSeek released such distilled models alongside R1.

A striking result from DeepSeek's report: R1-Zero, trained with RL alone and no SFT, learned on its own to produce longer reasoning and to re-check its work, with the response length growing during training. Its output was harder to read, which is why the released R1 added a cold-start SFT stage.

Input and output: training phase vs prediction phase

Training phase (RL)Prediction phase (use)
InputA problem plus a hidden reference answer or test suiteThe user's prompt (and optionally a reasoning-effort setting)
What the model producesSeveral sampled reasoning traces with final answersOne reasoning trace, then the final answer
What happens nextA checker scores each final answer; the model is updated toward higher-reward tracesThe answer (and maybe a summary of the reasoning) is returned
Who checks correctnessAn automatic verifierNobody, unless we add a verifier or tests

Note the last row: in training, wrong answers are caught by the checker. In use, the model's confident-sounding reasoning still needs our own checks (tests, validation) for high-stakes outputs.

When to use an LRM, and when to use a regular LLM

  • Use an LRM for multi-step maths and science, debugging and non-trivial coding, planning with constraints, careful analysis of long documents, and agent tasks where one wrong step ruins the result.
  • Use a regular LLM for chat, rewriting, translation, summaries, simple extraction and classification, and anything latency-sensitive such as autocomplete or voice.
  • Mix them: route by difficulty, or let a fast model handle routine steps and call a reasoning model for the hard ones. Many APIs expose a reasoning-effort knob (e.g. low / medium / high) so one model can do both.

The developer tools company Job A (labelling issues) goes to a regular LLM or even a small model: it is simple, high-volume and latency-sensitive. Job B (root-causing a failing test) goes to an LRM with a high reasoning effort, and its proposed fix is validated by actually running the tests.

Popular LRMs we should know

Notable reasoning models

  1. OpenAI o1: First widely used reasoning model; showed accuracy rising with thinking time.
  2. DeepSeek-R1: Open weights and a public recipe: RL with verifiable rewards (GRPO), plus distilled smaller models.
  3. Thinking modes everywhere: Anthropic Claude extended thinking, Google Gemini 2.5 thinking, OpenAI o3/o4-mini, Qwen3 thinking mode, gpt-oss reasoning effort levels.
  4. Reasoning as a dial: Models such as DeepSeek-V4 offer several modes (e.g. Non-think, Think High, Think Max) in one model.

Names and versions change quickly; the pattern to remember is that reasoning has become a setting on many general models rather than only a separate model family.

Common mistakes when using LRMs

  • Using them for everything. Simple tasks become slow and expensive, and models can overthink, talking themselves out of a correct easy answer.
  • Ignoring reasoning tokens in the budget. Thinking tokens are usually billed as output tokens and count against output limits even when hidden. Set a sensible effort level and max tokens.
  • Over-prompting. Adding "think step by step" or long few-shot examples is often unnecessary and can even hurt; vendors generally advise giving a clear goal, constraints and success criteria instead.
  • Trusting the trace as proof. A fluent reasoning trace can still end in a wrong answer. Verify with tests, calculations or sources.
  • Expecting unlimited scaling. Studies have found that on puzzles beyond a certain complexity, reasoning models can still collapse; more thinking is not a guarantee.

Pause and think: A team switches its customer FAQ bot from a regular LLM to an LRM. Answers barely improve, but latency triples and the bill doubles. What went wrong?

FAQ answers are simple lookups that do not need multi-step reasoning, so the extra thinking adds cost and latency without benefit. Use a regular LLM (plus retrieval) for the FAQ and reserve the LRM, or a higher reasoning effort, for genuinely hard queries.

Going one level deeper

We said GRPO compares each sampled answer with the average of its group. Let us do that by hand. For one problem the model samples 4 reasoning traces, a checker gives each final answer a reward of 1 or 0, and each trace gets an advantage: its reward minus the group average. (Implementations usually also divide by the spread of the group; we skip that to keep the numbers readable.)

Advantages for four sampled answers to one problem (illustrative groups)
Group rewardsGroup averageAdvantagesWhat the update does
1, 0, 0, 10.5+0.5, −0.5, −0.5, +0.5Makes the two correct traces more likely and the two wrong ones less likely
0, 0, 0, 10.25−0.25, −0.25, −0.25, +0.75A strong push toward the one trace that worked
1, 1, 1, 11.00, 0, 0, 0Nothing: the problem is already too easy
0, 0, 0, 00.00, 0, 0, 0Nothing: no trace shows what a good answer looks like

The last two rows explain a practical point about training data. A problem teaches the model only when some samples succeed and some fail. Problems that are far too easy or far too hard give an advantage of zero for every trace, so they cost compute and change nothing. Good training sets sit at the edge of what the model can currently do, and that edge moves as the model improves.

A second piece of arithmetic is the token bill. Suppose a visible answer is 200 tokens and the hidden reasoning before it is 3,000 tokens (illustrative). We are billed for 3,200 output tokens, 16 times the visible answer, and the same 3,200 count against the output limit. If the limit were 2,000, the model would run out during the reasoning and return no answer at all. An empty or cut-off reply from a reasoning model is very often this budget problem, not a model failure.

Practice: try it yourself

The earlier script spent test-time compute in parallel, by voting. Here we will simulate the sequential kind: a model that tries, checks its own answer, and backtracks when the check fails. We give it a larger and larger budget of tries and watch accuracy and tokens.

practice_check_and_retry.py

import random
random.seed(0)
P_SOLVE = 0.4          # chance one attempt is right (illustrative)
P_CATCH = 0.8          # chance the self-check notices a wrong attempt (illustrative)
TOKENS_PER_TRY = 500   # thinking tokens spent per attempt (illustrative)
def solve(max_tries):
# try, check, and backtrack until the check passes or the budget runs out
for attempt in range(1, max_tries + 1):
correct = random.random() < P_SOLVE
looks_ok = correct or random.random() > P_CATCH   # a wrong answer can slip by
if looks_ok or attempt == max_tries:
return correct, attempt
trials = 20_000
print("max tries  accuracy  avg thinking tokens")
for budget in [1, 2, 4, 8, 16]:
results = [solve(budget) for _ in range(trials)]
accuracy = sum(ok for ok, _ in results) / trials
tokens = sum(tries for _, tries in results) / trials * TOKENS_PER_TRY
print(f"{budget:>9}  {accuracy:>8.3f}  {tokens:>19.0f}")
# With endless tries, accuracy is capped by the wrong answers the check lets through
ceiling = P_SOLVE / (P_SOLVE + (1 - P_SOLVE) * (1 - P_CATCH))
print(f"ceiling with this checker: {ceiling:.3f}")

Output:

max tries  accuracy  avg thinking tokens
1     0.401                  500
2     0.598                  741
4     0.726                  906
8     0.767                  957
16     0.770                  954
ceiling with this checker: 0.769

Now change it:

  • Set P_CATCH = 1.0, a perfect self-check. Predict the ceiling and the accuracy at 16 tries before running.
  • Set P_CATCH = 0.0, a check that never notices errors. Predict the accuracy and the average tokens for every budget.
  • Set P_SOLVE = 0.1, a much harder problem, with P_CATCH back at 0.8. Predict the ceiling from the formula on line 25, then run it.

Pause and think: Accuracy stays near 0.77 even with 16 tries. Why does a bigger thinking budget not fix the remaining errors?

The limit is the check, not the budget. One in five wrong attempts passes the self-check and is returned as the answer, and once an answer passes, the model stops, so the extra tries are never used. Only a better way of verifying (tests, a calculator, a stricter checker) raises the ceiling.

Pause and think: The budget doubles from 8 to 16 tries, but average thinking tokens stay at about 955. Why?

The budget is a cap, not a spend. Each attempt passes the check with probability 0.4 + 0.6 × 0.2 = 0.52, so a run needs about 2 attempts on average and almost never reaches 8. Raising the cap changes almost nothing. The tiny difference between 957 and 954 is sampling noise.

Quick summary

  • An LRM is an LLM trained to think in a long reasoning trace before answering.
  • More test-time compute (longer thinking or more samples) buys accuracy on hard problems.
  • RL with automatically verifiable rewards is the core training ingredient; GRPO is a common method.
  • Use LRMs for hard, checkable, multi-step work; use regular LLMs for simple, fast tasks.
  • Budget for reasoning tokens, prompt simply, and verify outputs.

Key takeaways

  • An LRM is an LLM trained to reason in a long chain of thought before answering.
  • Test-time compute (longer thinking, more samples) is a new way to buy accuracy.
  • RL with automatically verifiable rewards (maths, code, puzzles) is the core training method.
  • LRMs are slower and cost more tokens, so use them for hard, multi-step, checkable problems.
  • Prompt simply, budget reasoning tokens, and verify answers instead of trusting the trace.

Key terms

  • Large Reasoning Model (LRM): An LLM trained to produce a long reasoning trace before its final answer.
  • Chain of thought: Intermediate reasoning steps written out before the final answer.
  • Test-time compute: Extra computation spent while answering, such as longer reasoning or more samples.
  • Self-consistency: Sampling several answers and choosing the most common one.
  • RLVR: Reinforcement learning with verifiable rewards: rewards come from automatic correctness checks.
  • GRPO: Group Relative Policy Optimization: an RL method that scores each sample against its group's average reward.
  • Reasoning effort: An API setting that controls how much a model thinks before answering.

← 7.1 Small Language Models: Big Capability in Compact Form · 7.3 Recursive Language Models: Self-Referential Generation →