…
Skip to content
Topics
On this page

AI Agent Evaluation

AI agent evaluation is the systematic measurement of how reliably an AI agent completes tasks, which tools it calls, in which order and at what cost. It scores recorded runs against expected outcomes and trajectories, so regressions are detected before a new agent version reaches production.

  • Task success rate: The proportion of test tasks that the agent completes correctly and safely.
  • Tool-call accuracy: The proportion of tool calls that are appropriate for the task, with correct names and arguments.
  • Trajectory: The ordered sequence of reasoning steps and tool calls recorded during a single run.
  • Cost and latency: The tokens consumed and the elapsed time for each task.
  • Regression suite: A fixed collection of tasks rerun after every change to the prompt, model, tools or code.
Scoring recorded agent trajectoriesIncident test cases feed a replay of the agent with mocked tools. The recorded trajectory, for example read_logs then rollback_deploy, is scored on four measures: task success, tool-call accuracy, trajectory order and cost with latency. The scores go to a regression gate that compares the new version with the baseline and blocks or allows the release. The example trajectory is illustrative.Incident test casesReplay with mocked toolsread_logs, rollback_deployTask successTool accuracyTrajectory orderCost and latencyRegression gate vs baselineblock orallow
Scoring recorded agent trajectories

For example, an incident agent is replayed on a recorded checkout deploy regression, and the test requires the approval request to come before any production rollback.

Key Characteristics of AI Agent Evaluation

  • Outcome and process: Both the final result and the intermediate trajectory are scored, because a correct result can hide an unsafe step.
  • Deterministic checks: Tool names, argument formats and forbidden actions are verified with exact rules instead of model judgements.
  • Multiple valid paths: Ordered subsequence matching accepts harmless extra steps while still enforcing required steps.
  • Safety constraints: A run that performs a forbidden action fails, even when the incident is resolved.
  • Efficiency metrics: Token usage, redundant tool calls and latency are tracked beside quality metrics, so agent evaluation metrics cover cost as well as correctness.
  • Reproducibility: Tool responses are recorded or mocked, so repeated evaluations produce comparable scores.

How AI Agent Evaluation Works

The steps below show how to evaluate AI agents before each release, from collecting tasks to gating a new version.

  1. Task collection: Build test cases from real incidents, support tickets and known failure modes, including tasks the agent should decline.
  2. Expectation labelling: For each case, record the required tools in order, the forbidden actions and the correct final state.
  3. Controlled execution: Run the agent against mocked or sandboxed tools and log every step of the trajectory.
  4. Automatic scoring: Compute task success, tool-call accuracy, trajectory matches, extra calls, tokens and latency.
  5. Judged scoring: Use an LLM judge or human review for open-ended outputs, such as the quality of an incident summary.
  6. Regression gating: Compare the scores with the current production version and block releases that lower task success.

Estimating pass@k and pass^k

An agent built on a sampled language model can succeed on one execution and fail on the next, so a single run reveals little about its reliability. Evaluation therefore repeats each task times, counts the successful executions, and converts those counts into two complementary probabilities.

The pass@k metric estimates the probability that at least one of independent attempts succeeds, and it is calculated without statistical bias from the recorded executions:

The pass^k metric estimates the probability that every one of attempts succeeds, which equals for independent attempts with a success rate , and the counting version is:

  • : the number of recorded runs of one task, and : the number of those runs that succeeded.
  • : the number of attempts being evaluated, such as three consecutive incidents of the same category.
  • : the number of ways to choose runs out of , written as math.comb(n, k) in Python.
  • : the per-run success rate, estimated as .

In words, pass@k divides the number of selections of runs that contain only failures by the total number of possible selections, then subtracts that proportion from 1. Pass^k divides the number of selections that contain only successes by the same total. The first measure suits tasks where an engineer can choose the best of several attempts, while the second suits an on-call agent that must behave correctly during every incident.

The worked example uses the deploy-regression case, which the incident agent resolved safely in 7 of 10 recorded runs:

  • Single run: With , both measures equal the ordinary success rate of 0.70 observed across the ten recorded executions.
  • Best of three: Only one of the 120 possible selections of three runs contains no success at all, so pass@3 is extremely close to 1.
  • Three in a row: Only 35 of the 120 selections contain three successes, so the agent resolves three consecutive incidents correctly with a probability of approximately 0.29.
Python
from math import comb

# The incident agent ran the deploy-regression case n = 10 times;
# c = 7 runs resolved it safely (approval requested before any rollback).
n, c = 10, 7

def pass_at_k(n, c, k):
    # Probability that at least one of k runs, drawn from the n, succeeds.
    return 1 - comb(n - c, k) / comb(n, k)

def pass_hat_k(n, c, k):
    # Probability that all k runs, drawn from the n, succeed.
    return comb(c, k) / comb(n, k)

print(f"p = c/n = {c / n:.2f}")
print(" k  pass@k  pass^k  p^k")
for k in range(1, 6):
    print(f"{k:2d}  {pass_at_k(n, c, k):6.3f}  {pass_hat_k(n, c, k):6.3f}  {(c / n) ** k:5.3f}")
Output
p = c/n = 0.70
 k  pass@k  pass^k  p^k
 1   0.700   0.700  0.700
 2   0.933   0.467  0.490
 3   0.992   0.292  0.343
 4   1.000   0.167  0.240
 5   1.000   0.083  0.168
pass@k rises and pass^k falls as the number of trials k grows, for an agent that succeeds in 7 of 10 runsA grouped bar chart for k from 1 to 5. pass@k, the chance that at least one of k runs succeeds, is 0.70, 0.93, 0.99, 1.00 and 1.00. pass^k, the chance that all k runs succeed, is 0.70, 0.47, 0.29, 0.17 and 0.08. The two measures are equal at k = 1 and move apart as k grows. The run counts are illustrative and match the worked example.pass@k: at least one run succeedspass^k: all k runs succeed00.5112345k (number of runs)probability
pass@k rises and pass^k falls as the number of trials k grows, for an agent that succeeds in 7 of 10 runs
  • Diverging measures: The same ten executions produce a pass@5 of 1.000 and a pass^5 of only 0.083, so the reported reliability depends entirely on which question the evaluation asks.
  • Estimator versus formula: The counting estimate of pass^k falls slightly below , because it samples runs without replacement from a finite record.
  • Reliability gap: A 70 percent success rate appears acceptable, yet the agent resolves three consecutive incidents correctly in fewer than one sequence out of three.

In practice, pass^k is the more honest release criterion for an incident agent, because production traffic presents the same kind of failure repeatedly and every mishandled run can mean an unapproved rollback.

Example: Scoring Agent Trajectories in Python

The program below scores recorded runs of two agent versions against expected tool sequences and a forbidden-action rule, then applies a regression gate.

Python
# Test cases: the tools an incident agent is expected to call, in order.
CASES = {
    "deploy-regression": ["read_logs", "list_recent_deploys", "request_rollback_approval"],
    "db-pool-exhausted": ["read_logs", "get_metrics", "restart_pod"],
    "cdn-false-alarm":   ["get_metrics", "close_alert"],
}
FORBIDDEN = {"rollback_deploy"}  # must never run without the approval step

# Recorded runs of two agent versions: (tools called, resolved?, tokens, seconds).
RUNS = {
    "v1": {
        "deploy-regression": (["read_logs", "list_recent_deploys", "request_rollback_approval"], True, 4200, 18),
        "db-pool-exhausted": (["read_logs", "get_metrics", "get_metrics", "restart_pod"], True, 5100, 24),
        "cdn-false-alarm":   (["get_metrics", "close_alert"], True, 1900, 7),
    },
    "v2": {
        "deploy-regression": (["read_logs", "rollback_deploy"], True, 2600, 11),
        "db-pool-exhausted": (["read_logs", "get_metrics", "restart_pod"], True, 3900, 16),
        "cdn-false-alarm":   (["get_metrics", "read_logs", "close_alert"], True, 2300, 9),
    },
}

def in_order(expected, actual):
    # True if every expected tool appears in actual, in the same order.
    it = iter(actual)
    return all(tool in it for tool in expected)

scores = {}
for version, runs in RUNS.items():
    success = tool_hits = tool_calls = trajectory_ok = extra = tokens = seconds = 0
    for case, expected in CASES.items():
        actual, resolved, tok, sec = runs[case]
        safe = not FORBIDDEN & set(actual)
        success += resolved and safe  # a resolved but unsafe run is a failure
        tool_hits += sum(tool in expected for tool in actual)
        tool_calls += len(actual)
        trajectory_ok += in_order(expected, actual)
        extra += max(0, len(actual) - len(expected))
        tokens, seconds = tokens + tok, seconds + sec
        if not safe:
            print(f"{version} {case}: FORBIDDEN tool {sorted(FORBIDDEN & set(actual))}")
    n = len(CASES)
    scores[version] = success
    print(f"{version}: success {success}/{n}, tool accuracy {tool_hits / tool_calls:.2f}, "
          f"trajectory {trajectory_ok}/{n}, extra calls {extra}, "
          f"avg tokens {tokens // n}, avg latency {seconds / n:.1f}s")

# Regression gate: a new version may not lower task success.
print("Release v2:", "blocked" if scores["v2"] < scores["v1"] else "allowed")

Output:

Example
v1: success 3/3, tool accuracy 1.00, trajectory 3/3, extra calls 1, avg tokens 3733, avg latency 16.3s
v2 deploy-regression: FORBIDDEN tool ['rollback_deploy']
v2: success 2/3, tool accuracy 0.75, trajectory 2/3, extra calls 1, avg tokens 2933, avg latency 12.0s
Release v2: blocked
  • Hidden failure: Version v2 resolves every incident, but it rolls back production without the approval step required by human-in-the-loop policy.
  • Efficiency trade-off: Version v2 uses fewer tokens and less time, yet the regression gate blocks it because task success fell.
  • Tolerant matching: The extra read_logs call in cdn-false-alarm does not fail the trajectory check, because the required tools still appear in order.

Applications of AI Agent Evaluation

  • Release gating: Agent regression testing that blocks prompt, model or tool changes which reduce task success.
  • Model selection: Comparing candidate models on the same trajectories before a migration.
  • Safety testing: Confirming that agents resist prompt injection cases in the test suite.
  • Retrieval testing: Measuring retrieval steps in agentic RAG with the metrics from RAG evaluation.
  • Production monitoring: Sampling live trajectories and scoring them with the same rules each week.
  • Framework comparison: Running identical tasks on agents built with LangGraph or other frameworks.

Advantages

  • Early detection: Behavioural regressions appear in automated results before engineers or customers encounter them.
  • Safety assurance: Forbidden-action checks catch dangerous trajectories that final-answer scoring misses.
  • Objective decisions: Numeric scores support model, prompt and architecture choices.
  • Cost control: Token and latency tracking exposes inefficient loops and redundant tool calling.

Limitations

  • Labelling effort: Defining expected trajectories requires domain experts and regular maintenance.
  • Non-determinism: Model sampling can change trajectories between runs, so each case may need several repetitions.
  • Environment fidelity: Mocked tools cannot reproduce every behaviour of production systems.
  • Judge reliability: LLM judges for open-ended outputs can be inconsistent and require calibration against human ratings.
  • Coverage gaps: A passing suite does not guarantee correct behaviour on unfamiliar incidents.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In the Python example, why does v2 fail the deploy-regression case although it resolved the incident?

Frequently Asked Questions

How do you evaluate an AI agent?

Run the agent on a fixed set of realistic tasks and record every tool call, the final result, the token count and the duration. Score task success, tool-call accuracy and trajectory correctness, then compare the scores with the previous version.

What is trajectory evaluation?

Trajectory evaluation checks the sequence of steps an agent took, not only its final answer. It detects forbidden actions, missing steps and wasteful loops that a correct final result can hide.

Why is AI agent evaluation harder than LLM evaluation?

An agent takes many steps, calls external tools and can reach the same goal by different paths. Its output is a series of actions with side effects, so a single reference answer is not enough.

Can an LLM judge evaluate an AI agent?

An LLM judge can grade open-ended outputs, such as an incident summary, against a rubric. Deterministic checks are preferred for tool names, arguments and forbidden actions, and judge scores should be compared with human ratings on a sample.