
Which LLM Evaluation Tools Produce Evidence, and Which Produce a Number
Author:
The LayerLens Team
Last updated:
Published:
Picture a procurement review six months after the purchase. The evaluation dashboard shows the support agent's pass rate climbing across four releases, and the slide reads as progress. Then someone from risk asks who wrote the criteria behind the number, and which record holds the reference answer the score was measured against.
Nobody in the room can answer either question. The tool works and the trend line moves, and what the team bought was a number with a login attached and nothing behind it.
Choosing among LLM evaluation tools is a decision about which failure modes the team will never see. The lists ranking for the phrase compare licenses, prices and metric counts, and none tells a buyer what a class of tool cannot prove.
TL;DR
The phrase LLM evaluation tools covers six distinct jobs, and most shortlists mix them.
Three questions separate tools that produce evidence from tools that produce a number: was there a reference answer, who authored the criteria, what artifact survives the run.
The MT-Bench human-agreement study, arXiv:2306.05685, Table 5, setup S2, found a GPT-4 judge agreeing with human experts 85% of the time, above the 81% agreement between the experts themselves. The same paper documents the biases that number hides.
LangChain's State of Agent Engineering report, from 1,340 responses, found 52.4% of teams running offline evals on test sets, 37.3% running online evals, and 59.8% of evaluating organizations still using human review.
A pass rate carries no coverage map. Suppose 5% failure per step across a ten-step task: completion lands at 59.9% before any other problem shows up.
The strongest single ask on a vendor call is one artifact from the last thirty days, tied to a trace the team already owns.
One Search Term Covers Six Different Jobs
Six jobs sit under this phrase, and each one defines done differently:
Tracing a single run: what the agent did, step by step
Offline scoring against a fixed set: how a known group of cases performed
Gating a change before merge: whether this version ships
Testing a prompt: which wording produced the better output
Evaluating retrieval: whether the right context reached the model
Guarding runtime behavior: whether production is drifting right now
A tracing tool answers what happened. An offline scorer answers whether the output matched a reference. A gate answers whether the change ships. Teams buy one of these expecting another, which is the most common purchase mistake in the category and the one no ranking page names.
A team that runs two of the six jobs will often describe itself as covering four, and the gap shows up when a score improves and nobody can name which class of failure moved with it.
Three Questions Decide Whether a Tool Can Prove Anything
What was the reference answer? A tool that scores against no fixed reference measures style. Correctness needs something to be correct against, such as an answer key that existed before the run started or a labeled set. Ask where the reference lives and whether the same one holds across agent versions.
Who authored the criteria? A default rubric encodes the tool vendor's idea of good. Ask for the name of the person who wrote the criteria the pass rate came from. A named author is how a rubric survives a review meeting, because a reviewer can ask that person what they meant.
What artifact records the check? The artifact is a record carrying the prompt version, the model pin, the judge version and a timestamp. Screenshots of dashboards fail this question. So does a chart that cannot name the judge that produced it.
Call the set the evidence triad: reference, author, artifact. Run it in order. The second question does most of the work, and it is the one buyers skip.
Judges are load-bearing in this category, so the agreement number earns its place here. The MT-Bench human-agreement study, arXiv:2306.05685, Table 5, setup S2, measured a GPT-4 judge against expert votes and found 85% agreement with human experts on non-tied comparisons, above the 81% agreement between the experts themselves. The same paper documents position bias, verbosity bias and self-enhancement bias. Those biases travel with the judge into any tool that ships a default, which is why the criteria author matters more than the metric count. LLM as a judge works through the mechanics of a default judge and the drift that shows up when the provider updates it.
The survey data says how far most teams have gotten. LangChain's State of Agent Engineering report, published June 2026 from 1,340 responses, found 52.4% of teams running offline evals on test sets and 37.3% running online evals, while 59.8% of evaluating organizations still use human review. Human review is the ground truth the other two calibrate against, and a tool that cannot show which human wrote the rubric has cut that calibration loop.
Where the Three Tool Classes Land on the Evidence Triad
Group the market by class. Three classes absorb almost every logo on a best-of list, and each class fails the triad in a different place.
Class | What it scores | Reference answer in the system | Who authors the criteria | Artifact that survives the run |
|---|---|---|---|---|
Offline frameworks | Model outputs against a dataset the team supplies | Usually present, because the dataset carries it | The team writes the metric, unless a built-in judge prompt ships one | Test output in CI logs, versioned with the repo |
Tracing-first tools | Runs as they happen, plus scores attached to a trace | Usually absent, since these record behavior | The team, through a scorer attached to a trace | The trace itself, retained per the retention setting |
Eval-workflow products | Whatever the team configures, plus vendor judge templates | Present when the dataset carries one | Often the vendor template by default | Experiment record inside the vendor's system |
[INSERT IMAGE: llm-evaluation-tools-evidence-triad.png - Branded data visualization: a three-by-three grid, rows for offline frameworks, tracing-first tools and eval-workflow products, columns for reference answer, criteria author and artifact, each cell marked "usually present", "usually absent" or "vendor default", with the one failing cell per row highlighted in the accent color. Alt text: Grid showing where offline frameworks, tracing-first tools and eval-workflow products each fail one of the three evidence questions.]
Read the table by column, because the failures do not overlap. Offline frameworks fail on the artifact when a team runs them locally and never stores the output; the score exists in a terminal scrollback and nowhere else. Tracing-first tools fail on the reference when the trace is the only record, because a trace shows behavior and a reviewer cannot tell from a trace alone whether the behavior was right. Eval-workflow products fail on the criteria author, because the default template decides what counts as good unless the team overrides it, and overriding a default takes a decision, a name and a version.
The classes also combine, which is what a production stack looks like after the first year. A team records runs with one tool and scores them with another, then reconciles the two by trace id, and that reconciliation is where the artifact question gets decided. Teams that already maintain an in-house LLM evaluation framework usually sit in the first class, and the metric definitions in the repo are the asset worth carrying forward.
A Pass Rate Carries No Coverage Map
Ninety-five percent on a set that never included the failure mode tells a buyer nothing about that failure mode. The number reassures while hiding the blind spot. LLM evaluation metrics lays out which slice of a run each metric covers, and none of them covers the slice the set left out.
Compound that across a multi-step agent and the arithmetic turns. Suppose 5% failure per step across a ten-step task. Completion lands at 59.9% before any other problem shows up, and compounding failure math for agents works through what that does to a per-step pass rate.
Coverage also fools the team that buys volume. Asking for 500 test cases produces 500 lexically varied cases with near-identical underlying behavior, which is the coverage illusion: volume wearing the costume of variety. The measured fix is trajectory diversity, the variance in the paths an agent takes across a set. A shortlist that cannot answer how it measures path variance will report a coverage number nobody can audit.
A coverage map answers the question with three columns: the failure class in the buyer's language (wrong tool called, hallucinated figure, refused a valid request, leaked context), the cases that exercise that class, and the evidence that each case ran. Ask which failure class the last release broke that the suite missed, because the answer names the coverage gap faster than any audit.
What the Evidence Triad Does Not Settle
It does not tell a team how many tools to run, and many production stacks end up with two: one recorder, one scorer. It does not repair vague criteria, and a precise number about a bad rubric stays bad. It says nothing about cost, latency, data residency or procurement rules, all of which can legitimately outrank evidence for a given buyer.
A regulated team with a residency constraint may be right to pick the weaker evaluator. That constraint is recoverable: when the rule changes, the team can move the data. An unproven pass rate stays unproven, because nothing in it says what was checked.
Run the Triad on a Shortlist This Afternoon
Put the three questions in columns and one candidate per row. The exercise takes about twenty minutes and the output is a sentence the reader can repeat to a vendor.
Candidate | Reference answer | Criteria author | Artifact |
|---|---|---|---|
(fill) | named, versioned, or absent | named person, or vendor default | record with prompt, model pin, judge version, timestamp |
The single strongest ask is one artifact from the last thirty days, tied to a trace the team already owns. A working answer names the criteria author and opens an artifact with the judge version on it. A deflection describes the dashboard, offers a demo, or explains that a success manager can pull something together. Teams weighing a from-scratch option against a purchase should read why building your own LLM evaluation framework goes wrong first, because the buy-versus-build answer usually follows from the criteria question.
[INSERT IMAGE: stratix-judge-detail-version-snapshot.png - Real Stratix screenshot of a system judge's detail view (for example the RAG Quality judge) showing the judge name, the version number, the plain-English evaluation goal and the pinned judge model, with a trace evaluation's stored judge snapshot open beside it. Alt text: Stratix judge detail view showing a judge name, version number, evaluation goal and pinned model next to the judge snapshot stored with a trace evaluation.]
The screen above is what the third question looks like when it has an answer. In Stratix, every judge carries a version number and a version history. Every trace evaluation stores a judge snapshot with the judge name, the version, the evaluation goal and the model used, so a verdict from March still says which judge produced it. [Marin question: does the judge record show who created or last edited a custom judge, so the second question has a field of its own?]
Frequently Asked Questions
What are LLM evaluation tools?
Software that scores a model's or an agent's outputs against criteria. The category covers six jobs: tracing runs, offline scoring against a fixed set, gating changes before merge, testing prompts, evaluating retrieval, and guarding runtime behavior. Most tools do two of the six well.
What is the difference between an LLM evaluation tool and an observability tool?
An observability tool records what happened in production. An evaluation tool decides whether what happened was acceptable, which requires a reference answer and criteria. Many products do both, and the evidence triad tells a buyer which half is real.
Do LLM evaluation tools need an LLM judge?
No. Exact match, regex, JSON-schema validity and semantic similarity are code graders that run without a model and reproduce exactly. A judge is needed only when the criteria require reading comprehension, and then the judge version becomes part of the record.
How accurate is LLM-as-a-judge?
The MT-Bench human-agreement study found a GPT-4 judge agreeing with human experts 85% of the time on non-tied comparisons, above the 81% agreement between the experts themselves. The same study documents position, verbosity and self-enhancement biases, so the agreement holds only while the rubric and the judge version stay fixed.
How many evaluation tools does a team need?
Many production stacks settle on two: one that records runs and one that scores them, reconciled by trace id. A single tool that claims all six jobs usually does two, and the other four end up in spreadsheets nobody versions.
What should a buyer ask a vendor before choosing an evaluation tool?
Three questions: what the reference answer was, who authored the criteria, and what artifact records the check. Then ask for one artifact from the last thirty days tied to one of the team's own traces.
Run the three questions against the current shortlist, then read the three jobs every evaluation product has to do for the unit-of-work question that follows. The judge library, judge versions and the snapshot stored with each trace evaluation are documented in the Stratix docs. Load twenty of the team's own traces and put the three questions to the result in Stratix.