Lesson 11.3 · 24 min
The Agent Loop: Observe, Think, Act, Repeat
Every AI agent, from a tiny script to a coding assistant, runs the same few lines of code over and over; what are they, and how do they know when to stop?
In short: The agent loop is the small piece of ordinary code that repeatedly calls the model, runs any tools it asks for, feeds the results back, and stops when the model gives a final answer or a limit is hit. Each pass is one think-act-observe cycle. Getting the loop right (stop conditions, error handling, parallel calls, budgets) is what separates a demo from a dependable agent.
The big picture
In the last two lessons we met AI agents and function calling. Function calling gives us one round trip: the model asks for a tool, we run it, we send back the result. But real tasks rarely finish in one round trip. A research question might need five searches; a bug fix might need reading three files, editing one, and running the tests twice.
The agent loop is what turns one round trip into as many as the task needs. It is surprisingly small: often 20–40 lines of code. Yet nearly every agent framework, under its abstractions, is running this same loop.
Think of it like cooking from a recipe you are improvising You taste the soup (observe), decide it needs salt (think), add a pinch (act), then taste again. You repeat until it tastes right, and you stop for other reasons too: the guests have arrived (time limit) or you have run out of salt (a tool keeps failing). The agent loop is that taste-adjust-taste cycle, written as code.
What is the AI agent loop?
The AI agent loop is a while loop around an LLM call. Each iteration (called a turn or step) does three things: call the model with everything so far, execute whatever tools it asked for, and append the results to the history. The loop ends when the model replies without asking for any tool, or when a safety limit triggers.
Two parts with very different natures work together here:
- The model is probabilistic. It reads the history and decides: call a tool, or answer.
- The loop code is deterministic. It never decides what to do for the task; it executes, records, and enforces rules (limits, permissions, timeouts).
This split matters. We cannot fully control what the model will say, but we can fully control the loop. Safety, cost control, and reliability are mostly built into the loop, not the prompt.
Why an AI agent needs a loop
Why not ask the model to plan everything and call all tools at once? Because the right next step often depends on the result of the previous one. Suppose a user says: “Why did yesterday's nightly job fail?” The agent must:
- List yesterday's job runs (it does not know the run id yet).
- Read the log of the failed run (needs the id from step 1).
- Notice a “disk full” error and check disk usage on that machine (only now does it know which machine).
None of these steps could be written in advance. Each one is chosen after seeing the last observation. A loop is the only structure that allows that: decide, act, look, decide again. It also allows recovery: if a tool returns an error, the next turn can try something else.
Pause and think: If every task needed exactly one known tool call, would we need a loop?
No. A single function-calling round trip (or even a fixed workflow) would do. The loop exists because the number and choice of steps depend on intermediate results, which we cannot know in advance.
The think-act-observe cycle
Each pass through the loop is one think-act-observe cycle:
Observation is what makes this different from a model just writing a long plan. The model gets grounded feedback from the real world after each action, so a mistake in step 2 can be noticed and corrected in step 3.
The loop step by step
What the loop code does on every run
- Initialise: Build the message list: system prompt (role, rules), the user's goal, and the tool definitions. Set counters: steps = 0, tokens used = 0.
- Call the model: Send the whole message list. The model is stateless, so it must see the full history every time.
- Inspect the reply: No tool calls? It is the final answer: return it. Tool calls present? Append the assistant message to the history and continue.
- Execute tools: For each call: validate arguments, check permissions, run with a timeout, and catch exceptions. Turn the result or error into text.
- Append observations: Add one tool-result message per call, linked by call id, in order.
- Enforce limits: Increase counters. If steps, time, tokens or cost exceed the budget, stop and return a partial answer or an explanation.
- Repeat: Go back to “Call the model”.
Because the full history is re-sent on every call, the input grows each turn. If each step adds about the same amount of text, the total input tokens processed grows roughly with the square of the number of steps. Prompt caching (covered in the inference module) reduces the price of the repeated prefix, but long loops are still expensive.
The loop in real code
Here is the shape of the loop against a real API, written with the Anthropic Python SDK. Other SDKs look very similar; only field names change. It needs an API key, so it has no output shown; the runnable version follows.
agent_loop_real_api.py (shape only)
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
def run_agent(goal, tools, run_tool, max_steps=10):
messages = [{"role": "user", "content": goal}]
for _ in range(max_steps):
resp = client.messages.create(model=MODEL, max_tokens=1024,
tools=tools, messages=messages)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use": # no tool needed: done
return "".join(b.text for b in resp.content if b.type == "text")
results = []
for block in resp.content: # may be several calls
if block.type == "tool_use":
try:
out = run_tool(block.name, block.input)
except Exception as e: # errors become observations
out = f"error: {e}"
results.append({"type": "tool_result",
"tool_use_id": block.id, "content": str(out)})
messages.append({"role": "user", "content": results})
return "Stopped: step limit reached."And here is a runnable, offline version that focuses on the loop's control logic. The “model” is a scripted list of turns, so we can see three runs end in three different ways.
loop_stop_conditions.py
# An agent loop with three stop conditions and parallel tool calls.
def lookup_price(item): return {"apple": 3, "bread": 5, "milk": 4}[item]
TOOLS = {"lookup_price": lookup_price}
def run(script, max_steps=4):
history, seen = [], set()
for step in range(1, max_steps + 1):
turn = script[min(step - 1, len(script) - 1)] # the "LLM" output
if "final" in turn: # stop 1: model is done
return f"done at step {step}: {turn['final']}"
results = []
for name, arg in turn["calls"]: # may be several (parallel)
key = (name, arg)
if key in seen: # stop 2: stuck repeating
return f"aborted at step {step}: repeated {name}({arg})"
seen.add(key)
results.append(TOOLS[name](arg))
history.append(results) # observations go back in
print(f" step {step}: {turn['calls']} -> {results}")
return f"stopped: hit max_steps={max_steps}" # stop 3: budget
good = [{"calls": [("lookup_price", "apple"), ("lookup_price", "bread"),
("lookup_price", "milk")]},
{"final": "total is 12"}]
stuck = [{"calls": [("lookup_price", "apple")]}] # repeats forever
wander = [{"calls": [("lookup_price", x)]} for x in ["apple", "bread", "milk", "apple"]]
for name, script in [("good", good), ("stuck", stuck), ("wander", wander)]:
print(name)
print(" ", run(script, max_steps=3 if name == "wander" else 4))Output:
good
step 1: [('lookup_price', 'apple'), ('lookup_price', 'bread'), ('lookup_price', 'milk')] -> [3, 5, 4]
done at step 2: total is 12
stuck
step 1: [('lookup_price', 'apple')] -> [3]
aborted at step 2: repeated lookup_price(apple)
wander
step 1: [('lookup_price', 'apple')] -> [3]
step 2: [('lookup_price', 'bread')] -> [5]
step 3: [('lookup_price', 'milk')] -> [4]
stopped: hit max_steps=3Pause and think: In the output, why did the “good” run finish in only 2 steps even though it needed three prices?
Because step 1 contained three parallel tool calls, all executed in the same turn. Step 2 was the final answer. Without parallel calls it would have needed 4 steps (3 lookups + 1 answer).
Parallel tool calls in one turn
When a model returns several tool calls in one reply, the loop should handle all of them before calling the model again. Two practical rules:
- Run independent calls concurrently (threads, async, or a worker pool). Three 1-second API calls then take about 1 second, not 3.
- Return every result, in the matching order and with the matching ids, in the next message. If one call fails, still send its error as its result; never drop it.
Careful with side effects Parallel is safe for reads (lookups, searches). For writes that depend on each other (create a folder, then a file inside it), running them at the same time can break things. Many APIs let you turn parallel tool calls off, or the loop can run calls one by one when a tool is marked as having side effects.
How the loop knows when to stop
A loop without good exits is the most common source of runaway cost. A well-built loop has several independent stop conditions:
A good loop also tells the model about its limits, for example: “You have 10 steps. If you cannot finish, summarise what you found.” When a budget is hit, return a helpful partial result rather than an empty failure.
Common loop failures
| Failure | Symptom | Fix |
|---|---|---|
| Infinite loop | Same search repeated with tiny variations. | Max steps, duplicate-call detection, tell the model what it already tried. |
| Lost tool results | API error about a tool call without a result. | Always append exactly one result per call id, even on error. |
| Crash on tool error | One bad argument kills the whole run. | Catch exceptions; return the error text as the observation. |
| Context overflow | Long runs exceed the context window or slow down. | Truncate or summarise old tool outputs; keep large data out of the prompt. |
| Premature finish | Model answers before verifying (e.g. says tests pass without running them). | Require evidence: a check tool, or a verification step before finishing. |
| Huge observations | A tool returns a 2 MB file into the prompt. | Cap result size; return summaries or pages with a pointer to the rest. |
The most expensive bug No step limit plus a tool that keeps erroring. The model retries forever, and each retry re-sends the growing history. Every production loop needs a hard cap on steps and cost, no matter how good the model is.
Quick summary. The agent loop is a simple while loop: call the model, run requested tools, append results, repeat. Each pass is a think-act-observe cycle. The model decides; the loop executes and enforces rules. Handle parallel calls together, never drop a result, turn errors into observations, and always pair the natural finish with hard budgets and stuck detection.
Worked example, step by step
The failure table says “truncate or summarise old tool outputs”. Let us do that by hand once, with small illustrative numbers, so the idea is concrete. Our loop has a rule: the history sent to the model must stay under 1,000 tokens.
After three steps the history holds the system prompt and goal (200 tokens), a log file from step 1 (600), a short lookup from step 2 (100), and a second log from step 3 (300). That is 1,200 tokens: over the limit before step 4 even starts.
| Option | What the history becomes | Tokens | What we lose |
|---|---|---|---|
| Drop the oldest result | 200 + 100 + 300 | 600 | Everything from step 1, including the clue the agent followed |
| Cut every result to 150 tokens | 200 + 150 + 100 + 150 | 600 | The tails of both logs, which may hold the key line |
| Summarise the oldest result | 200 + 50 + 100 + 300 | 650 | Detail from step 1, but its main finding is kept in 50 tokens |
How the loop applies the third option
- Measure before every model call: Count the tokens of the history. Here: 1,200, which is more than 1,000.
- Pick the oldest large observation: The step 1 log is the oldest and the biggest. Recent results are usually the ones the model still needs in full.
- Replace it, do not delete it: Swap its content for a short note such as “step 1: log showed disk full on node-7”. The tool call and its result message both stay, so the history is still consistent.
- Never touch the goal: The system prompt and the user's goal are kept word for word. An agent that loses its goal wanders.
- Measure again: Now 650 tokens. The loop calls the model for step 4.
No option is free. Dropping is cheapest but forgets. Cutting is simple but blind. Summarising keeps the most meaning but may need an extra model call to write the summary. Many loops start with a size cap on every result and add summarising only when runs get long.
Practice: try it yourself
We will write a loop that handles two things the earlier example skipped: a tool that fails, and observations that are too big. The scripted model asks for a file with a typo in its name, reads the error, and recovers. The loop caps every observation and watches a history budget. We count words as a rough stand-in for tokens.
practice_resilient_loop.py
# An agent loop that survives tool errors and keeps its history under a budget.
FILES = {"notes.txt": "disk full on node-7 " * 12, # a long file (48 words)
"node-7.log": "cleanup job disabled since Monday"}
def read_file(name):
return FILES[name] # raises KeyError for a missing file
def fake_model(history):
"""Scripted LLM: reacts to the latest observation in the history."""
last = history[-1] if history else ""
if not history:
return ("read_file", "node7.log") # a typo: file is missing
if last.startswith("error"):
return ("read_file", "notes.txt") # recover: try another file
if "node-7" in last and "cleanup" not in last:
return ("read_file", "node-7.log") # follow the clue
return ("final", "Disk filled up because the cleanup job is disabled.")
def words(history): # crude stand-in for a token count
return sum(len(h.split()) for h in history)
history, MAX_OBS_WORDS, BUDGET = [], 10, 40
for step in range(1, 7):
action, arg = fake_model(history)
if action == "final":
print(f"step {step}: FINAL {arg}")
break
try:
obs = read_file(arg)
except Exception as e: # an error becomes an observation
obs = f"error: no file {e}"
kept = obs.split()[:MAX_OBS_WORDS] # cap the size of every observation
obs = " ".join(kept) + (" ...[cut]" if len(obs.split()) > MAX_OBS_WORDS else "")
history.append(obs)
print(f"step {step}: read_file({arg!r}) -> {len(obs.split())} words, history={words(history)}")
if words(history) > BUDGET:
print("stopped: history budget used up")
breakOutput:
step 1: read_file('node7.log') -> 4 words, history=4
step 2: read_file('notes.txt') -> 11 words, history=15
step 3: read_file('node-7.log') -> 5 words, history=20
step 4: FINAL Disk filled up because the cleanup job is disabled.Now change it:
- Set
BUDGET = 12. Predict the step at which the loop stops and whether the agent reaches its final answer. - Set
MAX_OBS_WORDS = 3. Predict what the model sees fromnotes.txt, and which branch offake_modelit takes next. Does it still find the cause? - Remove the
try/exceptand callread_file(arg)directly. Predict the output of step 1. Which row of the loop-failure table is this?
Pause and think: At step 2 the file has 48 words but the output says 11 words. Where did the other words go, and why is the ...[cut] marker worth adding?
The loop kept only the first 10 words and added one marker word, giving 11. The marker tells the model the observation is incomplete. Without it, the model could treat a cut-off file as the whole file and draw a wrong conclusion, or never think to ask for the rest.
Pause and think: This loop reaches the right answer even though its first tool call failed. Which part deserves the credit: the model or the loop?
Both, in different roles. The loop caught the exception and turned it into text in the history instead of crashing. The model then read that text and chose a different file. If the loop had crashed, the model would never have had the chance; if the model ignored the error, the loop's care would not have helped.
Key takeaways
- The agent loop is a short while-loop: call the model, run requested tools, append results, repeat.
- Each pass is a think-act-observe cycle; observations let the agent react to real results.
- The model decides; the deterministic loop executes, validates, and enforces budgets.
- Always combine the natural finish with hard limits (steps, tokens, cost, time) and stuck detection.
- Turn errors into observations, return a result for every call, and cap the size of tool outputs.
Key terms
- Agent loop: The repeating code that calls the model, executes requested tools, and feeds results back until a stop condition.
- Turn (step): One iteration of the loop: one model call plus execution of any tool calls it returned.
- Think-act-observe: The cycle inside each turn: reason about the next move, take an action with a tool, read the result.
- Stop condition: A rule that ends the loop, such as a final answer, a step budget, or stuck detection.
- Stuck detection: Checks that notice the agent repeating itself or making no progress, and stop or redirect it.
- Budget: A hard limit on steps, tokens, money or time that guarantees the loop terminates.
← 11.2 Function Calling: Giving LLMs Tools to Act on the World · 11.4 ReAct Agents: Interleaving Reasoning and Acting →