Modern AI Engineering

Lesson 14.4 · 24 min

Agent Observability: Traces, Spans, and Debug Signals

A customer says our agent took 30 seconds and then refunded the wrong order: how do we find out exactly what it did, step by step?

In short: AI agent observability is the ability to see and understand what an agent did on every run: each model call, prompt, tool call, result, decision, token count, latency and cost. It is built from traces made of nested spans, plus metrics and logs. Observability lets us debug failures, control cost and latency, detect quality drift and feed real failures back into evaluation.

What is an AI agent, and what is observability?

An AI agent is a program in which a large language model (LLM) repeatedly decides what to do next: call a tool, look something up, ask the user, or finish. One user request can trigger many model calls and tool calls, and the path is chosen by the model at run time.

Observability is a term from software operations: a system is observable if we can understand what is happening inside it from the data it emits, without adding new code each time we have a question. That data is called telemetry. Monitoring watches known signals (“alert if error rate > 2%”); observability lets us investigate questions we did not anticipate (“why did this request take 30 seconds?”).

Think of it like an aircraft's flight recorder When something goes wrong in flight, investigators do not guess. They read the flight recorder: every control input, instrument reading and cockpit conversation, in order, with timestamps. Agent observability gives every agent run its own flight recorder, so we can replay what it saw, decided and did.

What is AI agent observability, and why do we need it?

AI agent observability means capturing, for every agent run, the full chain of events (inputs, prompts, model outputs, tool calls with arguments and results, retrieved documents, errors, timings, token counts and costs) in a structured, linked form, and making it searchable, visual and measurable.

  • Debugging. Agents fail in new ways: a wrong tool choice, a malformed argument, a loop, a hallucinated order ID. Without a record of each step we can only guess.
  • Non-determinism. The same input can take different paths. We need to see the path that actually happened.
  • Cost control. Token usage varies a lot per run. One runaway loop can cost more than a thousand normal requests.
  • Latency. Users feel the total time; traces show which step is slow.
  • Quality drift. Model updates, prompt edits or new data can quietly degrade answers. Scoring live traffic catches it.
  • Safety and audit. For actions such as refunds, we need to show who asked for what, which tool ran and with which arguments.

Our running example: the refund agent handling “refund order 1009”. A user complains it was slow. We will find out why from its trace.

How is AI agent observability different from traditional observability?

The three pillars of observability

PillarWhat it isAgent example
LogsTimestamped records of individual events, often as structured text“tool.issue_refund timed out after 2,300 ms”
MetricsNumbers aggregated over time, cheap to store and alert onp95 latency, tokens per request, daily cost, tool error rate
TracesThe end-to-end path of one request through the system, made of linked spansThe whole refund run: plan → get order → refund (timeout) → retry → reply

For agents, traces are the most important pillar, because an agent's behaviour is a sequence of linked decisions. Metrics are usually computed from traces (for example, total cost = sum of the cost of every LLM span), and logs are often attached to spans as events.

Traces and spans

A trace represents one request from start to finish and has a unique trace ID. A span is one unit of work inside the trace, such as one LLM call, one tool call or one retrieval. Each span has a name, a start time and an end time (so a duration), a parent span ID that places it in a tree, a status (ok or error), and attributes: key-value details like model name, token counts or tool arguments.

The root span covers the whole agent run; child spans nest beneath it. Viewed as a timeline (a “waterfall”), the trace shows what happened, in what order, and where the time went. The code below builds and prints a toy trace for our refund run.

toy_trace.py

spans = []
PRICE_IN, PRICE_OUT = 3e-6, 15e-6     # illustrative $ per token
def span(name, parent=None, ms=0, **attrs):
# A span = one timed unit of work, linked to its parent
s = {"id": f"s{len(spans) + 1}", "name": name, "parent": parent, "ms": ms, "attrs": attrs}
spans.append(s)
return s
# One trace for one user request (durations are made up but fixed)
root = span("agent.run", user="u_42", task="refund order 1009")
span("llm.call", root, 820, model="big-model", in_tok=1200, out_tok=90)
span("tool.get_order", root, 140, status="ok")
span("tool.issue_refund", root, 2300, status="timeout")
span("tool.issue_refund", root, 310, status="ok")          # retry
span("llm.call", root, 640, model="big-model", in_tok=1500, out_tok=60)
root["ms"] = sum(s["ms"] for s in spans if s["parent"] is root)
def show(s, depth=0):
# Print the trace as a tree, like an observability UI would
print(f"{'  ' * depth}{s['id']} {s['name']:<18}{s['ms']:>5} ms  {s['attrs']}")
for child in [c for c in spans if c["parent"] is s]:
show(child, depth + 1)
show(root)
llm = [s for s in spans if s["name"] == "llm.call"]
cost = sum(s["attrs"]["in_tok"] * PRICE_IN + s["attrs"]["out_tok"] * PRICE_OUT for s in llm)
errors = sum(s["attrs"].get("status") == "timeout" for s in spans)
slow = max(spans[1:], key=lambda s: s["ms"])
print(f"\ntotal={root['ms']} ms  llm_calls={len(llm)}  cost=${cost:.4f}  errors={errors}")
print(f"slowest: {slow['id']} {slow['name']} {slow['ms']} ms ({slow['attrs']['status']})")

Output:

s1 agent.run          4210 ms  {'user': 'u_42', 'task': 'refund order 1009'}
s2 llm.call            820 ms  {'model': 'big-model', 'in_tok': 1200, 'out_tok': 90}
s3 tool.get_order      140 ms  {'status': 'ok'}
s4 tool.issue_refund  2300 ms  {'status': 'timeout'}
s5 tool.issue_refund   310 ms  {'status': 'ok'}
s6 llm.call            640 ms  {'model': 'big-model', 'in_tok': 1500, 'out_tok': 60}
total=4210 ms  llm_calls=2  cost=$0.0103  errors=1
slowest: s4 tool.issue_refund 2300 ms (timeout)

Pause and think: Using the trace, what is the cheapest fix for the slow run?

Lower the timeout on issue_refund (or fix the refund service), since the timed-out call alone took 2,300 ms of 4,210 ms and the retry then succeeded in 310 ms. Switching to a faster model would save far less. Without the trace, many teams would have blamed the LLM.

What to observe inside an agent, and the key metrics

  • Inputs and outputs: the user request, the final answer, and each model call's prompt and completion.
  • Model details: model name and version, temperature and other settings, input and output token counts, cached tokens.
  • Tool calls: tool name, arguments, result (or a summary of it), status, duration, retries.
  • Retrieval: the query, which documents came back, their scores.
  • Decisions: the agent's plan or reasoning summary, which branch it took, guardrail verdicts.
  • Context: user or session ID, prompt version, app version, so we can group and compare runs.
  • Feedback and scores: thumbs up/down, LLM-judge scores attached to the trace.
Key metrics for AI agent observability
MetricWhy it matters
Latency: end-to-end, p50 / p95, time to first tokenUser experience; p95 shows the slow tail that averages hide
Tokens and cost per request and per userBudget control; catches runaway loops
Steps (LLM and tool calls) per runRising counts signal loops or confused planning
Error rate: model API errors, tool failures, timeouts, invalid outputsReliability of each dependency
Task success and user feedback rateWhether users actually got what they needed
Online quality scores (judge or heuristic)Detect drift in correctness, groundedness, tone
Guardrail triggers and policy violationsSafety, abuse and attack monitoring

How AI agent observability works

From code to dashboard

  1. Instrument: Add tracing to the agent code. Many SDKs and agent frameworks offer automatic instrumentation that wraps every LLM and tool call in a span; we add custom spans and attributes for our own steps.
  2. Propagate context: The trace ID is passed along so that every child span, even in other services, attaches to the same trace.
  3. Export: Spans are batched and sent asynchronously to a collector or backend (commonly using the OpenTelemetry protocol), so tracing does not slow the agent down.
  4. Store and index: The backend stores traces and makes them searchable by user, error, cost, latency or attribute.
  5. Visualize and alert: Dashboards show metrics over time; alerts fire on spikes in cost, latency or errors; engineers open individual traces to debug.
  6. Evaluate and improve: Sampled traces are scored by LLM judges or reviewed by people. Bad traces become new test cases in the evaluation suite.

Tools and frameworks, and observability vs evaluation

Examples (the field moves fast; check current docs)
CategoryExamples
Open standardsOpenTelemetry (traces, metrics, logs) with its generative-AI semantic conventions, which are still evolving; OpenInference and OpenLLMetry conventions for LLM spans
LLM / agent observability platformsLangfuse, LangSmith, Arize Phoenix, Helicone, Weights & Biases Weave, MLflow Tracing
General observability vendors with LLM featuresDatadog, New Relic, Grafana and others accept OpenTelemetry data and offer LLM views
Built into agent frameworksMany agent SDKs emit traces of model and tool calls out of the box

Challenges and best practices

Challenges Privacy: prompts and tool results can contain personal or confidential data, so traces need redaction, access control and retention limits. Volume and cost: full prompts for every call are large; teams sample or truncate. Silent failures: a wrong answer looks like a success unless we attach quality scores or feedback. Long, branching traces from multi-agent systems are hard to read. Standards are still settling, so attribute names differ between tools.

  • Trace from day one, including during development; it is the fastest debugger for agents.
  • Give every run a trace ID and log it with user-facing errors so support can find the trace.
  • Record versions of prompts, models and tools as attributes, so we can compare before and after a change.
  • Redact sensitive data before export, and set retention rules.
  • Alert on cost, latency, error rate and steps per run, not just on crashes.
  • Attach feedback and judge scores to traces to make silent failures visible.
  • Close the loop: turn bad traces into evaluation test cases.

Real-world use A team sees daily cost double overnight. Sorting traces by cost reveals an agent stuck re-calling a search tool after a schema change in its results. They fix the parser, add an alert on steps per run, and add that failing case to their evaluation suite.

Common mistakes and how to spot them

Collecting traces is the easy half. Reading the numbers built from them is where teams go wrong, and the most common mistake is trusting an average. A small example shows why. Ten refund runs finish; nine take 1.0 second and one, stuck on a slow tool, takes 11.0 seconds. The numbers are illustrative.

One slow run, three different stories

  1. The mean: (9 × 1.0 + 11.0) / 10 = 2.0 seconds. No user waited 2 seconds. The mean describes a run that never happened.
  2. The median (p50): Sort the ten values and take the middle: 1.0 second. This is the typical experience, and it hides the slow run completely.
  3. A high percentile: p95 asks: how long did the slowest 5% take? With only 10 runs, the one slow run is the slowest 10%, so p95 is pulled far above 1 second (the exact value depends on how the tool interpolates). Now the problem is visible.
  4. Grow the sample: With 100 runs and still one slow one, that run is the slowest 1%. It sits above p95, so p95 goes back to about 1 second. We need p99 or the maximum to see it.
  5. The rule: Report p50 for the typical user, p95 or p99 for the unlucky ones, and keep the maximum in view. Then open the trace of the worst run rather than staring at the dashboard.
A symptom on the dashboard, and where to look in the trace
What the metric showsLikely causeWhat to open in the trace
Mean latency up, p50 flatA few very slow runsSort traces by duration; look for one long span such as a tool timeout
Cost up, request count flatMore tokens per run: a longer prompt, or more stepsCompare input tokens and step count per run before and after the change
Steps per run creeping upThe agent repeats a tool call that keeps failing or returns nothing usefulLook for the same tool name with the same arguments several times in a row
Error rate flat, complaints upSilent failures: the run “succeeds” with a wrong answerTraces with low judge scores or negative feedback; read input, retrieved data and output
One user or tenant dominates costAn unusual input, or an automated callerGroup traces by user or session ID

Two instrumentation mistakes make all of this impossible: spans without a parent link, so a run cannot be reassembled, and traces without version attributes, so we cannot say which prompt or model produced a bad run.

Practice: try it yourself

Earlier we looked inside one trace. Now we stand one level up: we take summaries of 40 finished runs and compute what a dashboard would show, namely latency percentiles, cost and error rate. Then we add a simple alert rule that finds a run stuck in a loop.

practice_trace_metrics.py

import numpy as np
rng = np.random.default_rng(11)
# 40 finished agent runs, as a tracing backend would summarise them.
runs = []
for i in range(40):
steps = int(rng.integers(3, 7))                  # normal runs: 3 to 6 steps
runs.append({"id": f"run-{i:02d}", "steps": steps,
"ms": int(steps * rng.integers(300, 700)),
"tokens": int(steps * rng.integers(800, 1500)),
"error": bool(rng.random() < 0.05)})
# One run got stuck in a loop, calling the same tool again and again.
runs[17].update(steps=38, ms=41000, tokens=95000)
ms = np.array([r["ms"] for r in runs])
tokens = np.array([r["tokens"] for r in runs])
PRICE = 5e-6                                         # illustrative $ per token
print(f"latency  mean {ms.mean():.0f} ms | p50 {np.percentile(ms, 50):.0f} ms | p95 {np.percentile(ms, 95):.0f} ms | max {ms.max()} ms")
print(f"cost     total ${tokens.sum() * PRICE:.2f} | mean per run ${tokens.mean() * PRICE:.4f}")
print(f"errors   {sum(r['error'] for r in runs)} of {len(runs)} runs")
# Alert rule: flag any run with far more steps than a typical run.
typical = np.median([r["steps"] for r in runs])
for r in runs:
if r["steps"] > 3 * typical:
share = r["tokens"] / tokens.sum()
print(f"ALERT {r['id']}: {r['steps']} steps (typical {typical:.0f}), {share:.0%} of all tokens")

Output:

latency  mean 3344 ms | p50 2489 ms | p95 3571 ms | max 41000 ms
cost     total $1.49 | mean per run $0.0372
errors   3 of 40 runs
ALERT run-17: 38 steps (typical 5), 32% of all tokens

One run out of 40 used about a third of all tokens. Look at the latency line: p95 (3,571 ms) gives no hint of a 41-second run, because a single run in 40 lies above the 95th percentile. The maximum and the step alert are what expose it.

Now change it:

  • Add a second stuck run after line 13: runs[5].update(steps=30, ms=33000, tokens=70000). Predict: does p95 move now? Why does two runs out of 40 change it when one did not?
  • Make the alert too sensitive: on line 26 change 3 typical to 1.0 typical. Predict roughly how many alerts fire. Would we still read them?
  • Remove the runaway run by commenting out line 13. Predict how the mean compares with p50 once the outlier is gone.

Pause and think: The mean latency (3,344 ms) is higher than the p50 (2,489 ms) and close to the p95. What does that pattern alone tell us, before we look at any trace?

That the distribution has a long tail: a few runs are far slower than the rest and drag the mean up. In a healthy, symmetric set of runs the mean sits near the median. A mean well above p50 is a cue to sort traces by duration and open the slowest ones. Here that leads straight to run-17.

Pause and think: The alert names run-17 and says it took 38 steps. Is that enough to fix the problem?

No. The metric tells us that something looped and which run. Only the trace tells us why: which tool was repeated, with what arguments, what it returned each time, and what the model decided after each result. The fix might be a tool that returns an unclear error, a missing stop condition, or a step limit. We need the spans to choose between them.

Key takeaways

  • Agent observability records every model call, tool call, decision, token and millisecond of each run.
  • Traces made of nested spans are the core; metrics and logs complete the three pillars.
  • Agents fail silently, so attach quality scores and feedback, not just errors and latency.
  • Instrument, export, store, visualize, alert, and feed bad traces back into evaluation.
  • Handle privacy and volume with redaction, sampling and retention rules.

Key terms

  • Observability: The ability to understand a system's internal behaviour from the telemetry it emits.
  • Telemetry: Data a system emits about itself: logs, metrics and traces.
  • Trace: The end-to-end record of one request, made of linked spans and identified by a trace ID.
  • Span: One timed unit of work in a trace, with a parent, status and attributes.
  • Attribute: A key-value detail on a span, such as model name, token count or tool argument.
  • OpenTelemetry: An open standard and toolkit for producing and exporting traces, metrics and logs.
  • p95 latency: The latency below which 95% of requests complete; it reveals the slow tail.

← 14.3 Evaluating AI Agents: Metrics and Methods That Work · 15.1 LLM Guardrails: Filtering Inputs and Outputs for Safety →