Modern AI Engineering

Lesson 14.3 · 25 min

Evaluating AI Agents: Metrics and Methods That Work

Our refund agent gave the customer the right answer, but it called the wrong tool twice and nearly emailed someone else: did it pass?

In short: Evaluating an AI agent means judging not just its final answer but the whole multi-step process: whether the goal was reached in the real environment, which path it took, how it used tools, how well it planned, and what it cost. Because agents are non-deterministic and act on the world, we test them in sandboxed environments, run each task several times, and combine state checks, trajectory checks, LLM judges and human review.

What is an AI agent, and what is agent evaluation?

An AI agent is a system in which a large language model (LLM) works in a loop: it reads the goal and current state, decides on an action (often a tool call such as searching a database or calling an API), observes the result, and repeats until it decides the task is done. The sequence of thoughts, actions and observations in one run is called a trajectory (or trace).

AI agent evaluation is the practice of measuring how well an agent accomplishes tasks: did it reach the goal, did it get there in a sensible and safe way, and at what cost in time, tokens and money?

Think of it like evaluating a new employee, not grading an exam An exam checks the final answer. A manager evaluating a new hire also cares how the work got done: did they use the right systems, ask for approval before spending money, avoid breaking anything, and finish in reasonable time? A lucky correct outcome reached by a risky process is still a problem.

Our running example: a customer-support refund agent with tools search_orders, get_order, issue_refund and send_email. The task: “Refund my cracked blender, I ordered it last week.” The ideal path is search → get order → refund.

Why agent evaluation is needed, and how it differs from LLM evaluation

  • Agents act. A wrong tool call can refund the wrong order, delete a file or email a stranger. Mistakes have side effects, not just bad text.
  • Errors compound. A 10-step task with a 95% chance of getting each step right succeeds end to end only about 60% of the time (0.95¹⁰ ≈ 0.60).
  • Many paths can be valid. Two different tool sequences may both reach the goal, so we cannot simply compare to one expected answer.
  • Runs vary. The same agent on the same task may succeed on Monday and fail on Tuesday. One run tells us little.
  • Cost and latency vary per run. An agent that loops 40 times to finish a 3-step task is broken even if it eventually succeeds.

Outcome evaluation and trajectory evaluation

Outcome evaluation asks one question: was the goal achieved? The most reliable way is to check the final state of the environment with code, not the agent's own claim. For the refund agent: does the orders database now show a refund for the right order, for the right amount, and nothing else? For a coding agent: do the repository's tests pass? Outcome checks are objective and allow any valid path.

Trajectory evaluation looks at the path. We can compare it to a reference trajectory (exact match, same steps in any order, or the reference steps appearing as a subsequence), or grade it with a rubric: no unnecessary steps, no forbidden actions, asked for confirmation before risky actions, recovered sensibly from errors. Trajectory checks catch “right answer, wrong way” runs, but strict path matching can unfairly fail a valid alternative path.

Pause and think: Pause and predict: run t2 in the code below calls search_orders twice, then completes the refund correctly. Should a strict exact-path check fail it?

Strict matching does fail it, yet the outcome is correct and the extra call is harmless, just a little wasteful. This is why we usually pair an outcome check (pass) with softer trajectory metrics (one extra step, slightly higher cost) instead of failing every deviation from a reference path.

Tool use evaluation and planning evaluation

Tool use evaluation checks each tool call: was the right tool chosen (tool selection), were the arguments correct and well formed (argument accuracy), were needed calls made (recall) and unneeded ones avoided (precision), and did the agent handle tool errors well? Function-calling benchmarks such as the Berkeley Function Calling Leaderboard focus on exactly this.

Planning evaluation asks whether the agent broke the task into sensible steps and adapted when things changed. We can grade an explicit plan if the agent writes one, or infer planning quality from the trajectory: did it gather information before acting, did it avoid loops, did it re-plan after a failure instead of retrying the same broken call forever?

The four types side by side, for the refund task
TypeQuestionExample check
OutcomeWas the goal achieved?Database shows a refund on order 1009 only
TrajectoryWas the path sensible and safe?No send_email without a reason; ≤ 5 steps
Tool useRight tools, right arguments?issue_refund(order_id=1009, amount=89.00)
PlanningGood decomposition and adaptation?Looked up the order before refunding; re-planned after a timeout

Key metrics for AI agents

  • Task success rate: share of runs where the outcome check passes.
  • pass@k: probability that at least one of k attempts succeeds. Useful when a human or verifier can pick the good attempt.
  • pass^k (pass-hat-k, introduced with the τ-bench benchmark): probability that all k attempts succeed. Measures reliability, which matters when every customer gets one attempt.
  • Efficiency: steps per task, tokens, wall-clock latency, dollar cost.
  • Tool-call accuracy: precision and recall of tools, argument correctness, tool error rate.
  • Recovery rate: how often the agent recovers after a tool error.
  • Safety: rate of forbidden actions, policy violations, or actions taken without required confirmation.

agent_metrics.py

from math import comb
expected = ["search_orders", "get_order", "issue_refund"]
runs = {  # 4 trials of the SAME task: tool calls made, and was the goal reached?
"t1": (["search_orders", "get_order", "issue_refund"], True),
"t2": (["search_orders", "search_orders", "get_order", "issue_refund"], True),
"t3": (["get_order", "issue_refund"], False),          # guessed the order id
"t4": (["search_orders", "get_order", "send_email"], False),
}
def tool_scores(calls):
called, want = set(calls), set(expected)
precision = len(called & want) / len(called)   # were the calls relevant?
recall = len(called & want) / len(want)        # were the needed calls made?
exact = calls == expected                      # strict trajectory match
return precision, recall, exact
for name, (calls, ok) in runs.items():
p, r, exact = tool_scores(calls)
print(f"{name}: success={str(ok):5} steps={len(calls)} "
f"tool_P={p:.2f} tool_R={r:.2f} exact_path={exact}")
n = len(runs)
c = sum(ok for _, ok in runs.values())
for k in (1, 2, 3):
pass_at_k = 1 - comb(n - c, k) / comb(n, k)    # at least 1 of k tries succeeds
pass_hat_k = comb(c, k) / comb(n, k)           # all k tries succeed
print(f"k={k}: pass@k={pass_at_k:.2f}  pass^k={pass_hat_k:.2f}")

Output:

t1: success=True  steps=3 tool_P=1.00 tool_R=1.00 exact_path=True
t2: success=True  steps=4 tool_P=1.00 tool_R=1.00 exact_path=False
t3: success=False steps=2 tool_P=1.00 tool_R=0.67 exact_path=False
t4: success=False steps=3 tool_P=0.67 tool_R=0.67 exact_path=False
k=1: pass@k=0.50  pass^k=0.50
k=2: pass@k=0.83  pass^k=0.17
k=3: pass@k=1.00  pass^k=0.00

Agent benchmarks

Public agent benchmarks give a common environment, tasks and automatic success checks, so different agents can be compared.

Some widely used agent benchmarks

  1. WebArena: Realistic self-hosted websites (shopping, forums, code hosting); the agent completes tasks through a browser, checked by end state.
  2. GAIA: General-assistant questions that need web browsing, tools and multi-step reasoning, with short verifiable answers.
  3. SWE-bench: Real GitHub issues from Python repositories; success means the project's tests pass after the agent's patch. A human-validated subset, SWE-bench Verified, followed in 2024.
  4. AgentBench: A suite of environments (operating system, database, games, web) for comparing LLMs as agents.
  5. τ-bench and OSWorld: τ-bench: tool-using customer-service agents talking to a simulated user under policies, reporting pass^k. OSWorld: computer-use tasks in real desktop operating systems.
  6. Terminal-Bench and others: Command-line tasks in sandboxed terminals, plus many domain-specific agent benchmarks.

As with LLM benchmarks, scores can be affected by contamination, harness details and saturation. A benchmark tells us about general agent skill; our own task suite tells us whether our agent works.

Methods, frameworks and tools

How to evaluate an agent, step by step

  1. Build a task suite: Collect realistic tasks from real usage, including edge cases: missing info, ambiguous requests, tool failures, requests the policy forbids.
  2. Create a sandboxed environment: Fake or test versions of every tool and database, reset to a known state before each run, so the agent can act without real-world harm.
  3. Simulate the user if needed: For conversational agents, an LLM can play the customer with a scripted goal and personality.
  4. Run each task several times: Record the full trace: every message, tool call, argument, result, token count and timing.
  5. Grade with layered checks: Code checks the final state; code checks trajectory rules; an LLM judge grades conversation quality and policy compliance; humans review a sample.
  6. Track and compare: Report success rate, pass^k, cost and safety per version, and re-run the suite on every prompt, model or tool change.
Examples of tools teams use (capabilities change quickly; check current docs)
ToolWhat it is commonly used for
LangSmith, Langfuse, Arize Phoenix, BraintrustTracing agent runs, building datasets from traces, running evaluations and judge scorers
DeepEval, RagasOpen-source libraries with LLM-judge metrics (including RAG and agent-related metrics)
Inspect (UK AI Security Institute)Open-source framework for writing evaluations, including agentic tasks in sandboxes
OpenAI Evals, promptfooFrameworks for defining test cases and running evaluations across models and prompts

Challenges and best practices

Common traps Trusting the agent's own “Done!” message instead of checking the environment. Running each task once and reading noise as progress. Grading only exact paths and failing valid alternatives. Testing against live production systems. Ignoring cost: an agent that succeeds after 60 tool calls may be unusable. Letting test tasks leak into prompts or examples.

  • Check outcomes in the environment state, with code, whenever possible.
  • Run multiple trials and report pass^k for customer-facing reliability.
  • Grade the path with rules, not only with a reference path: forbidden actions, confirmations, step limits.
  • Measure cost and latency next to success.
  • Include failure injection: make tools time out or return errors to test recovery.
  • Turn production failures into new test tasks, using traces from observability tools.
  • Keep humans reviewing samples, especially for safety-critical actions.

Pause and think: Our agent has pass@3 = 0.95 but pass^3 = 0.40. It answers customers directly with no human picking the best attempt. Which number describes the customer experience better?

pass^3 (and pass^1). Customers get one attempt each; there is no one to choose the best of three. pass^k shows the agent is inconsistent, so we should work on reliability, not celebrate pass@3.

Going one level deeper

We said that errors compound. Let us turn that into a tool for deciding what to fix first. If each step succeeds with probability p and a task needs n steps with no second chances, the task succeeds with probability pⁿ.

Real agents can notice a failed step and retry. Suppose the agent recovers from half of its step failures. The effective per-step rate becomes 0.95 + 0.05 × 0.5 = 0.975, and a 10-step task goes from 0.95¹⁰ ≈ 0.60 to 0.975¹⁰ ≈ 0.78. Recovery is worth measuring because it moves the whole curve.

Finding the step to fix (illustrative numbers)

  1. Run the task 100 times: Our refund agent succeeds in 65 runs and fails in 35.
  2. Record where each failed run first went wrong: From the trajectories: search_orders 2 runs, get_order 3, issue_refund 25, send_email 5. That adds up to 35.
  3. Read the pattern: One step causes 25 of the 35 failures. The agent is not “65% good” everywhere. It is very reliable at three steps and weak at one.
  4. Estimate the gain before doing the work: If a clearer tool description cuts issue_refund failures from 25 to 5, success rises from 65 to about 85 runs. Halving the failures of the other three steps together would gain only about 5.
  5. Re-run and re-count: After the fix, repeat the 100 runs. A new weakest step will appear. Agent improvement is this loop, repeated.

This is why trajectory data matters even when the outcome check is the final judge. The outcome tells us how often the agent fails. The first failing step tells us where to spend the next day of work.

Practice: try it yourself

We will write a small layered grader for the refund task. It checks three things in order of importance: the outcome in the database, a safety rule on refund size, and a step budget. Four recorded runs go through it, each with a different story.

practice_agent_grader.py

# A tiny layered grader for the refund task: "refund order 1009, 40 dollars".
REFUND_LIMIT = 50
MAX_STEPS = 5
runs = {   # each run: the tool calls made, then the final database state
"r1": ([("get_order", 1009), ("issue_refund", 1009, 40), ("send_email", 1009)],
{"refunded": {1009: 40}}),
"r2": ([("get_order", 1009), ("send_email", 1009)],
{"refunded": {}}),                       # said "done", did nothing
"r3": ([("get_order", 1009), ("issue_refund", 1009, 400), ("send_email", 1009)],
{"refunded": {1009: 400}}),              # refunded 10x too much
"r4": ([("search_orders", "helmet")] * 4 + [("get_order", 1009),
("issue_refund", 1009, 40), ("send_email", 1009)],
{"refunded": {1009: 40}}),               # correct but wasteful
}
def grade(calls, state):
outcome = state["refunded"] == {1009: 40}       # check the world, not the words
refunds = [c for c in calls if c[0] == "issue_refund"]
safe = all(c[2] <= REFUND_LIMIT for c in refunds)
efficient = len(calls) <= MAX_STEPS
return outcome, safe, efficient
print("run  outcome  safe   efficient  steps  verdict")
passed = 0
for name, (calls, state) in runs.items():
outcome, safe, efficient = grade(calls, state)
ok = outcome and safe                           # efficiency is only a warning
passed += ok
verdict = "PASS" if ok and efficient else "PASS (slow)" if ok else "FAIL"
print(f"{name}   {outcome!s:<7}  {safe!s:<5}  {efficient!s:<9}  {len(calls):>5}  {verdict}")
print(f"task success rate: {passed}/{len(runs)} = {passed / len(runs):.2f}")

Output:

run  outcome  safe   efficient  steps  verdict
r1   True     True   True           3  PASS
r2   False    True   True           2  FAIL
r3   False    False  True           3  FAIL
r4   True     True   False          7  PASS (slow)
task success rate: 2/4 = 0.50

Run r2 sent a confirmation email without refunding anything. A grader that trusted the email would have passed it. Run r3 did refund, but ten times too much.

Now change it:

  • Make the grader trust actions instead of state: replace line 18 with outcome = ("send_email", 1009) in calls. Predict which runs’ outcome value flips, and which run now passes that should not.
  • Add a run that refunds the wrong order: "r5": ([("get_order", 1010), ("issue_refund", 1010, 40)], {"refunded": {1010: 40}}). Predict its outcome and safe values. What does the safety rule fail to notice?
  • Loosen the step budget: set MAX_STEPS = 10 on line 3. Predict what changes in the table and what does not change in the success rate.

Pause and think: Run r3 already fails the outcome check. Why keep a separate safety check at all?

Because the two can come apart, and they are not equally serious. An agent could refund 400, notice, and correct it to 40: the final state passes, yet a forbidden action happened on the way. Or it could email another customer’s details and still complete the refund. Safety rules look at the actions taken, outcome checks look at the end state. We also want reports to separate “did not finish” from “did something it must never do”.

Pause and think: Run r4 reached the goal but called search_orders four times first. Our grader passes it with a warning. When should that become a hard failure?

When the extra steps carry a real cost or risk: a strict latency or cost budget, tools that charge per call or have side effects, or repeats that suggest the agent was stuck in a loop and escaped by luck. For read-only searches, a warning plus a tracked “steps per task” metric is usually right. Failing harmless detours would punish valid alternative paths.

Key takeaways

  • Agent evaluation judges the outcome in the environment, the trajectory, tool use, planning, cost and safety.
  • Check outcomes by inspecting the final state with code, not by trusting the agent's message.
  • Run tasks several times: pass@k measures “can it ever”, pass^k measures “does it always”.
  • Use sandboxed environments, simulated users, layered graders and failure injection.
  • Public benchmarks compare agents in general; our own task suite decides if our agent is ready.

Key terms

  • AI agent: An LLM-driven system that repeatedly chooses actions, often tool calls, observes results and continues until a goal is reached.
  • Trajectory: The full sequence of reasoning, actions and observations in one agent run.
  • Outcome evaluation: Checking whether the agent achieved the goal, ideally by inspecting the final environment state.
  • Tool-call accuracy: How correctly an agent chooses tools and fills in their arguments.
  • pass@k: Probability that at least one of k attempts at a task succeeds.
  • pass^k: Probability that all k attempts at a task succeed; a measure of reliability.
  • Sandbox: An isolated, resettable environment where an agent can act without real-world side effects.

← 14.2 LLM-as-Judge: Automating Evaluation with Another Model · 14.4 Agent Observability: Traces, Spans, and Debug Signals →