
AI Agent Evaluation Metrics: The 8 That Should Stop a Release
Author:
The LayerLens Team
Last updated:
Published:

An agent that books meetings returns “Booked for Tuesday at 2 p.m.” on 47 of 50 test tasks, so the dashboard shows 94%. Read the runs behind that number and 6 of the 47 booked over an existing meeting, 4 sent the invite twice after a timeout, and 9 took more than 15 steps on a task that needs 4. The 94% measured how often the agent said it finished, and a team that ships on it finds the other numbers in its support queue.
An agent evaluation metric is worth tracking when a bad value would stop a release. That test cuts most metric lists down to eight.
TL;DR
Agent metrics answer three questions: did the agent finish the job, did it get there safely, and would it do it again at a cost worth paying. A single success rate answers only the first, and only if it is checked against a reference instead of the agent’s own claim.
Task success should be graded against a reference answer or the end state of the system, never against the agent’s final message, because agents report success on tasks they did not complete.
Path metrics (correct tool use, forbidden actions, instruction following) catch the runs that end with the right answer by a route that breaks on the next task.
Repeat-run pass rate, steps per pass and cost per pass decide whether a passing agent is ready to run at volume.
Judge agreement with human reviewers is the metric behind the metrics. An unchecked judge puts an unknown error bar on every other metric.
Three Questions, Eight Metrics

Did it finish the job?
1. Task success against a reference. The share of tasks where the agent’s answer matches a reference answer written before the run. Without one, a judge ends up grading whether the answer sounds right, and confident wrong answers pass.
2. End-state correctness. For agents that change things (book, refund, update, file), the answer is the state of the system after the run. Check the calendar, the ticket, the record. The meeting-booking agent above said “Booked” on 47 tasks, and the calendar showed 6 of those 47 as double-booked.
Did it get there safely?
3. Tool-call correctness. The share of tool calls with the right tool, the right arguments and a correct reading of the result. This catches the agent that searched with the wrong customer ID, got lucky because the IDs shared a prefix, and passed. Code checks the parts that can be stated exactly (a required call happened, the ID matches); a judge reads the rest.
4. Forbidden actions. A count of runs that did something they must never do: a destructive call without confirmation, a write outside scope, a repeated identical call past a limit, a spend over budget. Track the raw count and read every run behind it. One refund to the wrong account is a release blocker no matter how many tasks passed.
5. Instruction following. The share of runs that obeyed the constraints in the system prompt and the task: ask before cancelling, stay under the fare cap, never quote a price the policy table does not list. These failures often sit next to a correct answer, which is why they need their own column.
Would it do it again at a cost worth paying?
6. Repeat-run pass rate. Run each task several times and count the tasks that passed every time. An agent that passes a task 3 times out of 5 will fail it about 2 times in 5 in production, even though a single run would have scored it a pass.
7. Steps per passing task. The median number of steps on tasks the agent passed, compared with the fewest steps the task needs. The meeting agent’s 9 runs over 15 steps on a 4-step task were passes that a customer would experience as a slow, confused assistant, and they are usually the first runs to fail when an API gets slower.
8. Cost per passing task. Total spend on the run divided by the number of passes. Dividing by tasks flatters an agent that fails cheaply, so divide by passes to see what a correct result costs. That is the number to compare when choosing between two foundation models for the same agent.
The Metric Behind the Metrics: Judge Agreement
Metrics 1, 3 and 5 usually come from an LLM judge reading the run. Before trusting any of them, take 30 to 50 runs, have a person grade them, and measure how often the judge agrees. A judge that agrees with people 70% of the time puts a wide error bar on every number it produces, and a change of a few points between two agent versions can sit entirely inside it.
Two habits keep judge numbers honest. Pin the judge version so every agent version is graded by the same judge, and re-check agreement whenever the judge’s instructions or underlying model change. LLM judge reliability covers agreement scores and the biases worth testing for.
Metrics That Look Useful and Rarely Change a Decision
Some numbers show up on every agent dashboard and almost never decide a release.
Average judge score on a 1 to 10 scale. A move from 7.2 to 7.4 has no threshold anyone ships on, so it never settles an argument about a release.
Final-answer similarity to a reference. Text overlap scores reward answers that sound like the reference and miss answers that are right in different words. For agents, check the fact or the end state directly.
Latency per step. Useful for infrastructure, misleading for quality. Total time to a passing result is the number users feel.
Tokens per run. Cost per pass already includes it, and in a form that accounts for failures.
Reading the Metrics Together
No single metric says an agent is ready. A version that raised task success from 80% to 86% and also added one forbidden action where the old version had none is a worse release. A version that held task success flat and cut cost per pass by 40% may be the better release.
The way to read them is by version. Hold the tasks and judges fixed, run the old and new agent, and compare each metric task by task, with regressions listed first. Step-level evaluation goes deeper on grading individual steps inside a run, and pinned traces covers keeping the test set fixed across versions.

Grading the Path and Comparing Versions in LayerLens
Judges for the path. Natural-language judges read a full trace and return pass or fail with a score and step-by-step reasoning. Write one per question: task done, tools used correctly, instructions followed. Each judge is versioned, and each evaluation stores the judge snapshot that produced it.
Every trace judged as it arrives. Automation rules attach a judge to traces matching an agent, framework, status or tag, with a sample rate and a daily cap, so production runs get the same metrics as test runs.
References from your own work. A custom benchmark built from a brief, documents or selected traces gives each task a reference answer, and its test cases stay editable until the first evaluation and are restricted after, so every version is scored on the same test.
Version comparison. Run comparison labels every task as regressed, improved or unchanged between two evaluations, with both answers and the reference side by side.

Frequently Asked Questions
What is the most important AI agent evaluation metric? Task success checked against a reference answer or the system’s end state. Every other metric explains why that number moved or whether it will hold at volume. Task success judged from the agent’s own final message is the least reliable number on most dashboards.
How is agent evaluation different from LLM evaluation? An LLM evaluation grades one output for one input. An agent evaluation grades a sequence of decisions, including tool calls and their results, and often a change to an outside system. That adds path metrics (tool use, forbidden actions) and end-state checks that single-output evaluation does not need.
How many times should each task be repeated? Three to five runs per task is enough to separate tasks the agent passes reliably from tasks it passes by chance. Use more on tasks that involve money or customer data, where a 1-in-5 failure is too often.
Should forbidden actions be a rate or a count? A count, reviewed run by run. A rate of 0.5% sounds small and still means one customer in two hundred had a refund sent to the wrong account. Any non-zero count on a high-impact action should block release until someone has read the runs.
Can these metrics be computed without a reference answer? Path metrics, repeat-run pass rate, steps and cost can, but task success needs one. A judge can still grade plausibility without a reference, but plausibility is the property that lets confident wrong answers through, so build references for at least the tasks that matter most.
How often should the test set change? Keep it fixed while comparing versions, and refresh it every four to six weeks so it keeps matching what users ask. When it changes, re-run the current version first to set a new baseline.
Pick Your Three Blockers
Before the next release, choose the three metrics that would stop it: usually end-state correctness, forbidden actions, and repeat-run pass rate on your highest-risk tasks. Grade the current version on them first, so the next version has something to beat.