Modern AI Engineering

Lesson 12.2 · 26 min

Loop Engineering: Designing Reliable Agentic Loops

Why does an agent that is clever on every single step still sometimes run in circles for fifty steps, or stop after two and proudly announce a job it never did?

In short: An AI agent works by running a loop: look at the situation, pick an action, run it, look at the result, and repeat. Loop engineering is the practice of designing that repeating cycle on purpose: what the model sees each turn, how actions run, how progress is tracked, and above all when and how the loop stops. Most agent failures are loop failures, and a handful of techniques (budgets, verification, loop detection, context trimming, recovery) fix most of them.

What is loop engineering? Loop + engineering

Split the name in two. A loop is a cycle that runs again and again: in an AI agent, one turn of the loop is “the model looks at everything so far, chooses an action, the action runs, the result is added to what the model will see next”. Engineering means designing something deliberately, with limits, measurements and failure handling, instead of hoping it works.

So loop engineering is the deliberate design of the cycle an agent repeats until a task is truly finished. It is one part of the bigger harness (the previous lesson): the harness is all the code around the model; the loop is the beating heart of that code.

Think of a cook tasting a soup A good cook does not add salt once and serve. They taste, adjust, taste again, and stop when it is right, or when dinner time arrives and they serve the best version so far. A bad cook either never stops adjusting, or serves without tasting. Loop engineering is teaching the agent to taste, to adjust sensibly, and to know when to stop.

Our running example: a coding agent asked to “make the failing test in utils.py pass”. Each turn it may run tests, read a file, edit a file, or declare that it is done.

Why do we need loop engineering?

A single model call is like a single move in a game. Real tasks need many moves, and each move depends on what the previous one revealed. That is why agents loop. But loops multiply risk: if each step has a small chance of going wrong, a long chain of steps has a large chance that something goes wrong.

A simple illustrative calculation: if each step is right 95% of the time and steps were independent, a 20-step task would be fully right only about 0.95²⁰ ≈ 36% of the time. Real steps are not independent, but the lesson holds: without checks and recovery, long loops decay. Loop engineering exists to stop small errors from compounding.

The second reason is cost. Every turn re-sends the growing conversation to the model. A loop that runs 3× longer than needed costs far more than 3×, because later turns carry more tokens. A loop with no stop rule can burn a budget overnight.

What is a loop in an AI agent?

Almost every modern agent, from coding assistants to research agents, runs the same basic cycle. It is often called the agent loop or the ReAct pattern (reason + act).

One turn of the loop

  1. Observe: Build the input for this turn: the goal, the instructions, the history of actions and results so far, and any fresh information.
  2. Decide: Call the model. It returns either a tool call (an action with arguments) or a final answer.
  3. Act: The harness runs the requested tool: run tests, read a file, search, call an API.
  4. Record: The result (output, or an error) is added to the history, so the next turn can see it.
  5. Check: Should we stop? The task may be verified done, a budget may be used up, or the agent may be stuck. If none of these, start the next turn.

The simplest loop and its problems

Here is the loop most people write first:

naive_loop.py (do not ship this)

history = [goal]
while True:
reply = model(history)
if reply.is_final_answer:
break                      # stop when the MODEL says it is done
result = run_tool(reply.tool, reply.args)
history.append(result)         # history grows forever

It looks fine and works in demos. In real use it has four hidden problems:

  • No upper bound. while True with no step or cost limit can run forever if the model never says “final answer”.
  • The model decides when it is done. Models sometimes declare success too early (“I fixed it!”) without running the tests. The stop rule trusts the very thing we are unsure about.
  • History grows without limit. After many turns the history no longer fits in the context window, or important early facts get buried and ignored.
  • No handling of errors or stuck states. A tool crash ends the program, and an agent that keeps repeating the same failing action is never noticed.

Pause and think: In the naive loop, which single line is responsible for the “declares success without doing the work” failure?

if reply.is_final_answer: break. The loop stops whenever the model claims to be done. A better rule stops only when an independent check (for example, the test suite) confirms the goal is met.

The parts of the loop that we must engineer

Every turn has a few decision points. Each is a place where we choose behaviour on purpose:

Design decisions inside the loop, with the coding-agent example.
PartQuestion we answerCoding agent example
Turn inputWhat does the model see this turn?Goal, last test output, files read so far, a short progress note
Action spaceWhich actions are allowed?run_tests, read, edit, say_done; no delete
Tool executionHow are actions run and limited?Sandbox, 60 s timeout, output cut to 4,000 characters
Observation formatHow are results shown back?Only the failing test names and the first error, not 2,000 lines of logs
Progress trackingHow do we know we are moving forward?Number of failing tests per turn; a to-do list the agent updates
Stop conditionsWhen does the loop end?Tests pass (verified), or 30 steps, or $1 spent, or stuck
RecoveryWhat happens after an error or a stuck state?Feed the error back once; after 2 repeats, nudge or escalate to a human

Three kinds of stop A healthy loop has a success stop (an external check says done), a budget stop (steps, time, tokens or money run out) and a stuck stop (no progress or repeated actions). Most broken loops are missing at least one of the three.

Prompt vs context vs loop engineering

Common ways a loop breaks

  • Infinite loop: the stop condition is never met. Often because “done” is fuzzy or the model never says it.
  • Repetition: the agent repeats the same action (re-reading the same file, re-running the same failing command) because nothing tells it that it already tried.
  • Premature success: the agent declares done without checking, or after a check that does not really test the goal.
  • Context overflow: history grows until it exceeds the window or the important parts get lost in the middle.
  • Error spiral: one tool error leads to a confused fix, which causes another error, and so on.
  • Goal drift: after many turns the agent starts solving a different problem (refactoring code nobody asked about).
  • Cost runaway: each turn is affordable, but the run takes 200 turns.

The most expensive mistake Letting the model be the only judge of “done”. It is the root of both premature success and, in the other direction, endless polishing. Always pair the model's claim with an outside check whenever one exists.

Techniques of loop engineering

Fixes, matched to the failures above

  1. Budgets: Hard caps on steps, wall-clock time, tokens and money. When a cap hits, stop and report the best result so far rather than crash. This bounds infinite loops and cost runaways.
  2. Verified stopping: Done means an external check passed: tests green, schema valid, API returned success. The model can propose done; the checker confirms it. This fixes premature success.
  3. Loop and repetition detection: Keep a short memory of recent actions. If the same action with the same arguments appears N times, intervene: tell the model it already tried that, change strategy, or stop.
  4. Context management: Summarise or drop old turns, keep tool outputs short, and keep a running progress note (goal, what is done, what is next) at the top. This fixes overflow and drift.
  5. Error recovery: Show errors to the model in a readable form, retry transient failures automatically, and after a few failed attempts escalate to a human instead of spiralling.
  6. Checkpoints and logs: Save state after each turn so a long run can resume after a crash, and log every turn so we can see where loops go wrong.

Pause and think: An agent re-runs npm test five times in a row with the same failing result. Which two techniques address this most directly?

Repetition detection (notice the identical action and intervene) and error recovery (after a few failures, change approach or escalate). A step budget would also stop it eventually, but only after wasting turns.

A complete example

Below is a small but complete engineered loop for our coding agent. The “model” is scripted so the run is repeatable, and it deliberately makes two classic mistakes: running tests before doing anything, and repeating the same edit. Watch how the loop handles both.

engineered_loop.py

# A scripted "model" that proposes one action per turn for: make tests pass.
script = ["run_tests", "read utils.py", "edit utils.py", "edit utils.py",
"edit utils.py", "run_tests", "say done"]
def fake_model(turn):
return script[turn] if turn < len(script) else "say done"
def tests_pass(state):
return state["edits"] >= 1          # the real check: one correct edit fixes it
def run_loop(max_steps=6, max_repeats=2):
state, history = {"edits": 0}, []
for step in range(max_steps):
action = fake_model(step)
# Guard 1: stop a repeating action before it burns budget
if history[-max_repeats:] == [action] * max_repeats:
print(f"step {step}: {action!r} repeated, nudging model")
history.append("nudge"); continue
history.append(action)
if action.startswith("edit"):
state["edits"] += 1
print(f"step {step}: {action}")
# Guard 2: done means the checker says done, not the model
if action == "say done" or action == "run_tests":
if tests_pass(state):
return f"finished in {step + 1} steps (tests pass)"
print("         tests fail -> keep going")
return f"stopped: step budget of {max_steps} used up"
print(run_loop())

Output:

step 0: run_tests
tests fail -> keep going
step 1: read utils.py
step 2: edit utils.py
step 3: edit utils.py
step 4: 'edit utils.py' repeated, nudging model
step 5: run_tests
finished in 6 steps (tests pass)

In six steps the loop survived an early test run (it simply kept going), blocked a third identical edit, and stopped only when the checker confirmed success. Change max_steps to 4 and the same script ends with “step budget used up”, which is also a correct, safe outcome.

Worked example, step by step

Earlier we said that a loop running 3× longer costs far more than 3×. Let us check that with small numbers. Suppose the first turn sends 500 tokens (goal, instructions, tools), and each turn adds 300 tokens of action and result to the history. All numbers are illustrative.

Adding up the tokens

  1. Tokens per turn: Turn 1 sends 500, turn 2 sends 800, turn 3 sends 1,100. Turn n sends 500 + 300 × (n − 1).
  2. A 5-turn run: 500 + 800 + 1,100 + 1,400 + 1,700 = 5,500 tokens.
  3. A 15-turn run: The turns grow from 500 up to 4,700. Their sum is 39,000 tokens.
  4. Compare: 3× the turns, but 39,000 ÷ 5,500 ≈ 7.1× the tokens. The cost of a run grows roughly with the square of its length.
  5. Now cap the history: If we trim or summarize so that no turn sends more than 1,400 tokens, the 15-turn run costs 19,200 tokens: about half.

Two lessons follow. A step budget alone is a weak cost limit, because late steps cost much more than early ones; a token budget measures what we really pay. And context management is not only about quality: it turns a curve that bends upward into a straight line.

Practice: try it yourself

We will build a loop with all three stops: a success stop from a checker, a budget stop measured in tokens, and a stuck stop that fires when the number of failing tests has not improved for two turns. The model's results are scripted as failing-test counts.

practice_three_stops.py

# Three stops in one loop: verified success, token budget and no progress.
def run(failing_per_turn, token_budget=6000, patience=2):
history, spent, best, stale = 500, 0, None, 0
for turn, failing in enumerate(failing_per_turn, 1):
spent += history                   # each turn re-sends the whole history
history += 300                     # and the new result makes it longer
if best is None or failing < best:
best, stale = failing, 0       # progress: fewer failing tests
else:
stale += 1                     # no progress this turn
print(f"  turn {turn}: failing={failing} spent={spent} stale={stale}")
if failing == 0:
return "success stop: the checker reports 0 failing tests"
if stale >= patience:
return f"stuck stop: no progress for {patience} turns (best={best})"
if spent + history > token_budget:
return f"budget stop: the next turn would pass {token_budget} tokens"
return "script ended"
# Failing-test counts that a model's edits might produce, turn by turn
for name, script in [("steady", [5, 3, 3, 1, 0]),
("stuck", [5, 4, 4, 4, 2, 0]),
("slow", [5, 4, 3, 2, 2, 1, 1, 0])]:
print(name)
print(" ", run(script))

Output:

steady
turn 1: failing=5 spent=500 stale=0
turn 2: failing=3 spent=1300 stale=0
turn 3: failing=3 spent=2400 stale=1
turn 4: failing=1 spent=3800 stale=0
turn 5: failing=0 spent=5500 stale=0
success stop: the checker reports 0 failing tests
stuck
turn 1: failing=5 spent=500 stale=0
turn 2: failing=4 spent=1300 stale=0
turn 3: failing=4 spent=2400 stale=1
turn 4: failing=4 spent=3800 stale=2
stuck stop: no progress for 2 turns (best=4)
slow
turn 1: failing=5 spent=500 stale=0
turn 2: failing=4 spent=1300 stale=0
turn 3: failing=3 spent=2400 stale=0
turn 4: failing=2 spent=3800 stale=0
turn 5: failing=2 spent=5500 stale=1
budget stop: the next turn would pass 6000 tokens

Now change it:

  • Set patience=1. Predict which of the three runs end differently, and at which turn.
  • Set token_budget=20000. Predict how the slow run ends and its final spent.
  • Change history += 300 to history += 0, as if we held the context at a fixed size. Predict spent at turn 5, and how the slow run ends now.

Pause and think: In the stuck script the counts are 5, 4, 4, 4, 2, 0, so the tests would have passed at turn 6. The loop stopped at turn 4. Was that a mistake?

It is a trade-off, not a bug. At turn 4 the loop knows only that two turns in a row brought no progress; it cannot see the future. With patience=2 we accept that some slow but working runs get cut, in return for never paying for runs that are truly stuck. If two-turn plateaus are normal for our task, we raise the patience. A good stuck stop also hands over the best state so far, so the work is not lost.

Pause and think: The budget stop fires when spent + history > token_budget, before the next turn runs, and not after spent has passed the budget. Why check ahead?

Because the cost of the next turn is already known: it will re-send the whole history. Checking ahead means we never overshoot, and we stop while there is still a clean state to report. Checking afterwards would let the most expensive turn of the run, the last one, go over the limit.

Where it works well and where it fails

Loop engineering shines when progress is checkable: code with tests, data pipelines with validation, form filling with a schema, search tasks with a clear target. The checker gives the loop a reliable signal for “keep going” vs “stop”.

It struggles when done is a matter of taste: “write a great marketing email”, “design a nice logo”. Without a reliable check, the loop either stops at the first plausible draft or polishes forever. Here we can still add an LLM judge or a rubric, but that check is itself fuzzy, so we should add a human review step.

When not to loop at all: if one model call can do the job (translate this paragraph), a loop only adds cost and latency. And when the steps are known in advance and branch in fixed ways, a predefined graph of steps is often clearer than a free-form loop; that is the topic of the next lesson.

In real products Coding agents such as Claude Code and Cursor's agent run exactly this kind of loop: they read, edit and run tests or commands, feed results back, and are bounded by permissions, limits and user interruption. Research agents loop over search and reading with a budget on the number of searches.

Key takeaways

  • An agent is a model in a loop: observe, decide, act, record, check.
  • Small per-step error rates compound over long loops, so loops need checks and recovery.
  • A healthy loop has three stops: verified success, budget exhausted, and stuck.
  • Never let the model be the only judge of done when an external check exists.
  • Repetition detection, context trimming and progress notes fix most long-run failures.
  • Loops work best where progress is checkable; fuzzy goals need human review.

Key terms

  • Agent loop: The repeating cycle in which a model picks an action, the harness runs it, and the result is fed back for the next decision.
  • Loop engineering: Deliberately designing an agent's loop: turn inputs, actions, progress tracking, recovery and stop conditions.
  • Stop condition: A rule that ends the loop: verified success, an exhausted budget, or detection that the agent is stuck.
  • Verified stopping: Ending the loop only when an independent check, not the model's claim, confirms the goal is met.
  • Budget: A hard limit on steps, time, tokens or money for one run.
  • Goal drift: When an agent, after many turns, starts working on something other than the original task.

← 12.1 Harness Engineering: The Scaffolding Around AI Agents · 12.3 Graph Engineering: Stateful Workflows for Agents →