
The world your agent was graded in
Author:
The LayerLens Team
Last updated:
Published:
AI evaluation is midway through a change in what a test even is. Most model benchmarks have been lists of questions: the model answered and a grader counted. An agent benchmark works differently, because an agent has to do things: talk to a customer, click through a website, run commands in a terminal. So the benchmark builds a test world around it, meaning everything in the run that is not the agent itself. The customer it talks to is another LLM playing a role. The website it works in is usually a saved copy, and the APIs it calls are wired to answer every time. Somebody built each of those pieces and chose how hard it pushes back, which is why the score measures the agent and the test world together.
TL;DR
Carnegie Mellon ran 451 real people and 31 LLM user simulators through the same 165 airline and retail support tasks: agents finished 63.6% with humans and 77.8% under the top general-purpose simulators.
The simulators were nicer than the people they replaced: GPT-4o's simulated customer wrote polite turns 49.0% of the time, real customers 15.3%.
Ohio State's live-web benchmark (300 tasks, 136 real websites, human-graded) put OpenAI's Operator at 61.3%; a simple search baseline fell from 51% on WebVoyager's frozen sites to 22% on live ones in 100-task samples.
Arizona State's sim-to-real frame names 4 places a test world can run easier than production: what the agent reads, which tools it gets, how the world answers, and what the score counts. Moving a tool-calling test from English to Chinese raised one model's error rate from 5.5% to 46.5%.
The Stratix record shows the spread inside one model: o4 Mini High scored 100% on AIME 2024 and 18.75% on Terminal-Bench (Terminus-1).
Before trusting an agent score, ask who played the user, how fresh the world was and what was allowed to fail in it, and whether the number would survive production settings.
The same model, 81 points apart
The Stratix record shows the spread. o4 Mini High scored 100% on AIME 2024, 30 competition math problems. The same model scored 18.75% on Terminal-Bench (Terminus-1), 80 tasks that hand it a working terminal: set up a git server, untangle Python dependencies, configure a web server. The 2 scores sit 81 points apart, and the recorded failures involved command syntax, dependency conflicts, and environment setups left unfinished at timeout, the work of operating the world.
[INSERT IMAGE: stratix-evaluation-results.png - Stratix evaluation results view showing per-prompt outcomes for a recorded run]
Image URL: https://i.imgur.com/XZMiy44.png
When a test world simplifies reality, the simplification can show up in the score as agent skill. One question runs through this piece: what world produced the score?
[INSERT IMAGE: te06-table.png - In the test world vs in production: the user, the state, the responses, the bill]
Image URL: https://mcusercontent.com/82dd3f8e78273bcc7c72a2f8c/images/eb359a99-c2c4-b346-4dc4-0794ac7d123a.png
The scripted customer is too nice
The first piece is who plays the customer. Benchmarks in the tau-bench family score an agent on support conversations, and the customer in those conversations is an LLM reading a persona card. The design runs through the harder tiers too: Sierra's tau2-bench has the simulated user operating tools alongside the agent in its dual-control telecom tasks, and Meta's Gaia2 wraps agents in a fully simulated app environment. Carnegie Mellon hired 451 real people to run the same 165 airline and retail tasks, and ran 31 LLM user simulators through the identical setup. With real people, agents finished 63.6% of tasks. Under the top general-purpose simulators, 77.8%. Those simulators were nicer than the people they stood in for. GPT-4o's customer wrote polite turns 49.0% of the time; real customers, 15.3%. A real person gets terse and holds back details. The top simulators explain themselves patiently and hand over complete answers. That cooperation inflates the score.
[INSERT IMAGE: te06-chart-simulated-users.png - Agent task success: 77.8% under top general-purpose LLM user simulators vs 63.6% with 451 real people]
Image URL: https://mcusercontent.com/82dd3f8e78273bcc7c72a2f8c/images/1c368f6f-3635-aa13-eafe-a3711d93b58d.png
The live web grades harder
The websites are the second piece. Web agents have reported success rates near 90% on WebVoyager, a test covering 15 sites with automatic grading that Ohio State found unreliable. The same group built a live-web benchmark: 300 tasks, 136 real websites, each run graded by 2 or more human annotators. The best agent, OpenAI's Operator, finished 61.3%. In 100-task samples, a simple search baseline scored 51% on WebVoyager and 22% on the live tasks. 29 of its 51 points did not survive the move to real websites.
[INSERT IMAGE: te06-chart-live-web.png - Search baseline: 51% on WebVoyager vs 22% on live-web tasks, 100-task samples]
Image URL: https://mcusercontent.com/82dd3f8e78273bcc7c72a2f8c/images/380ab489-0e76-596e-4391-d9e573fb1f35.png
The 4 settings on a test world
Arizona State put names on the settings, borrowing the sim-to-real frame from robotics. A test world can run easier than production in 4 places: what the agent reads, which tools it gets, how the world answers, and what the score counts. Language is the loudest of the 4. Run the same tool-calling test in Chinese and Qwen3-Next-80B's error rate climbs from 5.5% to 46.5%. Tool lists come curated in a test, while a real deployment piles up near-duplicates. Test APIs are free to skip the timeouts and partial failures production serves. And the score counts accuracy, with latency and cost left off the bill. A benchmark can sit at the easy end of all 4 at once and still get quoted as if it measured production.
Built worlds have a purpose. A controlled world keeps a test repeatable, and rerunning hundreds of live-web tasks on every commit is not practical. The problem is reporting the score without the test conditions.
So when the next benchmark number shows up in a launch review, ask 3 questions. Who played the user? How fresh was the world, and what was allowed to fail in it? Would the number survive the settings production runs on? If nobody can answer them, treat the score as unfinished.
Frequently Asked Questions
Why do LLM user simulators inflate agent scores?
The Carnegie Mellon study found the top general-purpose simulators explain themselves patiently and hand over complete information, where real customers were terse and withheld details. Politeness alone shows the gap: 49.0% polite turns for GPT-4o's simulated customer against 15.3% for real people. An agent that never has to ask a follow-up question looks more capable than it is.
Does the 14-point gap mean tau-bench scores are wrong?
No. The scores are correct for the world they were produced in. The gap means a tau-bench number is an upper bound on what the same agent does with real customers, and it should be reported with the simulator that produced it.
Why was WebVoyager's automatic grading unreliable?
Ohio State found the automatic judge disagreed with human graders often enough that reported success rates near 90% did not hold up. Their replacement benchmark, Online-Mind2Web, uses 2 or more human annotators per run on 136 live websites, where the best agent finished 61.3%.
What are the 4 sim-to-real settings, in plain terms?
Observation is what the agent reads: language, formatting, page state. Action is which tools it gets and how clean the list is. Transition is how the world answers, instant success in a test against timeouts and partial failures in production. Reward is what the score counts: accuracy alone, or accuracy plus latency and cost.
Is a controlled test world a bad thing?
No. A controlled world is what makes a test repeatable, and rerunning hundreds of live-web tasks on every commit is not practical. The failure is quoting the score without the conditions. Record who played the user, how fresh the world was, what was allowed to fail, and what the score counted, and the number becomes usable.
How does Stratix record the test conditions?
Every Stratix run stores the model version and endpoint, the benchmark and task set, the harness, and per-prompt results, so a score can be read next to the setup that produced it. The o4 Mini High spread above (100% on AIME 2024, 18.75% on Terminal-Bench) comes from those per-prompt records.
Sources
Mind the Sim2Real Gap in User Simulation for Agentic Tasks (Carnegie Mellon). 451 humans, 165 tasks, 31 LLM user simulators; top general-purpose simulators lifted agent success to 77.8% against a 63.6% human baseline.
An Illusion of Progress? Assessing the Current State of Web Agents (Ohio State). Online-Mind2Web: 300 live-web tasks across 136 websites, human-graded; OpenAI's Operator topped the field at 61.3%.
The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective (Arizona State). Decomposes test-world divergence into observation, action, transition, and reward; a tool-calling test moved from English to Chinese raised error rates by 41 points.
tau2-bench (Sierra) and Gaia2 (Meta), the advanced tier of simulated-user and simulated-app agent benchmarks.
This article is adapted from Issue 06 of Trace Evidence, the weekly evaluation newsletter from LayerLens. Run your own benchmarks at stratix.layerlens.ai.