Your Eval Judge Changed. The Dashboard Did Not Tell You.

Author:

The LayerLens Team

Last updated:

Published:

The same judge prompt scored at 87% agreement with human labels one week and 79% the next. Nothing changed in the prompt or the evaluation setup. The provider updated the model behind the judge name, and the evaluation team caught it because they tracked judge version explicitly. According to the LayerLens GEPA guide, most teams do not track this.

The 87-to-79 figure is a documented example from a published evaluation guide. The mechanism behind it operates on every team that uses an LLM as a judge and references the judge by a model name without a frozen version hash. OpenAI, Anthropic, Google, and Mistral have all shipped version changes behind stable API names. A prompt that scored well on the prior model version scores worse on the current one. The dashboard shows the number going down but not why.

This piece teaches you how to detect when your judge changed, how to distinguish judge drift from a real regression, and how to set up a drift detection check that runs in milliseconds with no extra LLM calls. The method is platform-independent. It works on Stratix or on any open-source eval pipeline.

The Hidden Fork in Every Eval Comparison

Every version-over-version comparison in an evaluation pipeline assumes the measurement instrument is stable. The judge that scored version N is the same instrument that scores version N+1. When the judge model changes between the two runs, that assumption breaks.

The score delta between version N and version N+1 has three components. The agent change is the thing the team wants to measure. The judge change happens when the provider updated the model, so the judge scores differently and the score moves. The interaction between the two means a prompt that worked on the old judge behavior may not work on the new one, producing a score delta that reflects neither independently.

The dashboard reports one number. The evaluation tooling reports pass rates and averages. Nothing in that output tells the team which component moved.

A team comparing last week's eval run to this week's sees a pass rate that dropped from 82% to 76%. The natural inference is that the Tuesday deployment caused a regression. The engineer who shipped that deployment spends two hours investigating and finds nothing wrong. The agent is fine. The judge changed.

The arrow from deployment to score change does not run straight. It forks. One path runs through the agent. The other runs through the provider model update. The dashboard shows the arrow but not the fork.

Score movement has four possible sources: system change, judge model update, judge prompt change, and noise. Each produces the same dashboard artifact: a score change. Score alone cannot tell you which one happened.The Compiler That Changed Silently

A software team would not deploy a code change against a CI suite where the test runner output changed because the compiler was updated between runs. The test runner and the compiler are both supposed to be stable, and change control governs both. Agent evaluation accepts a different arrangement by default.

The judge model is the test runner. The provider is the compiler vendor. When Anthropic ships Claude Opus 4.8 behind the same endpoint name that delivered Claude Opus 4.7, it does what GCC does when it ships a new patch release. The difference is that GCC updates the version string. The model endpoint does not.

A LangChain survey from mid-2026 found that 52% of teams run offline evaluations and 37% run online evaluations. It did not measure how many of those evaluations reference the judge by a frozen version identifier versus a provider-managed latest endpoint. The gap is the number of teams that cannot tell their judge version has changed.

Provider model updates happen. The update is invisible to the tooling. The eval pipeline stores the judge name but not the model hash. A comparison between last month's eval run and this month's run that shows a 6-point pass-rate delta might reflect a genuine improvement in the agent or a change in the judge scoring behavior. It could also reflect a shift in which traces made the population, or all three bound together in a signal no one can separate.

How the Anchor-Set Mechanism Works

The fix is cheap and decisive. It requires no extra LLM calls or provider changelogs.

Select 15 to 20 traces that represent the kinds of tasks your agent handles in production. Label them. The labels are ground truth: correct, wrong, or partially correct. The labeling effort matters less than the persistence. Any fixed set is better than a moving window.

Run the judge against the anchor set at baseline and record the kappa agreement against the human labels. Run the judge against the anchor set again at every comparison point. Compute kappa against the same baseline.

The rule is two-way. If kappa fell, the judge changed and the score delta between versions may reflect drift. Do not trust the comparison until the judge stabilizes. If kappa held but the live eval score moved, the agent changed and the score delta reflects a real behavioral shift.

The anchor set costs nothing to maintain. Fifteen traces, five minutes of labeling, one kappa computation per comparison window. It runs in CI in milliseconds because the traces are stored locally and the judge scores them against frozen labels. No extra LLM calls. No dependencies on provider changelogs.

This mechanism answers whether a score movement is real, the result of judge drift, or just noise. No vendor changelog required.Reproducing Judge Drift on Identical Traces

The demonstration is straightforward. Take a stored evaluation log. Re-grade it with two model versions of the same judge name on identical inputs. The verdicts disagree.

The same trace, the same judge prompt, two model versions, different scores. That is the core claim in one experiment. Any team with stored eval logs and access to two model versions can reproduce it in an afternoon.

Kappa agreement is the natural drift metric for this comparison. A second signal worth tracking is grade-parse failure rate: a model version change can produce a different output format the grader cannot parse. The rate of unreadable verdicts is itself a drift indicator.

Open-source drift-detection tooling has emerged across the ecosystem in 2026. Projects like DriftWatch, evidently's data drift monitors, and several smaller experiment-tracking extensions each address a piece of this problem from a different angle. The fact that multiple independent teams built tooling for judge stability confirms that drift is a production pain teams recognize only once they have felt it themselves.

The Three-Stage Detect-Recalibrate Workflow

Detection is the first stage. The anchor set tells you drift happened but does not fix it.

The full workflow has three stages.

Stage 1: Anchor set detects drift. Run the frozen traces against the judge. Compute kappa against baseline. If kappa fell, drift is confirmed.

Stage 2: Re-graded logs confirm the source. Re-run stored evaluation logs with the prior judge version and the current judge version on the same traces. This distinguishes a provider-side model update from a judge configuration change or a natural distribution shift.

Stage 3: Recalibration restores the instrument. The LayerLens GEPA workflow re-optimizes the judge prompt against the current distribution. This is a shipped Stratix capability documented in the published GEPA guide. It re-establishes agreement with human labels on the judge's current behavior.

Each stage produces a known state. Stage 1 leaves the team knowing whether drift exists. Stage 2 reveals what caused it. Stage 3 delivers a recalibrated judge. None of the stages requires unreleased features or platform-specific tooling.

The distinction matters because detection and remediation involve different investments. An anchor set is a five-minute setup cost. A re-grading workflow requires stored evaluation logs. GEPA recalibration requires the Stratix platform or a parallel prompt optimization pipeline. Teams that have not confirmed drift exists should not invest in downstream stages until Stage 1 returns a signal.What Stratix Captures Today and What Is Missing

Stratix already stores the data needed for a stronger version of the anchor-set check. The TraceEvaluation record carries a judge_snapshot field with four values: name, version, evaluation_goal, and model_name. When a team pulls two eval runs and compares the judge_snapshot values, any difference in version or model_name confirms the judge identity changed between the runs.

The SDK documentation states that you can "reuse the same judge across different product areas, version it, and track how its scores change over time." Stratix acknowledges judge versioning as a tracking dimension. The data model was built for this use case.

Three open engineering questions define the gap between the data Stratix stores and the feature a user can act on.

First: does judge_snapshot expose through the public SDK today, or is it internal-model only? If the field is internal, the snapshot check is not actionable for users. The anchor-set method described here is platform-independent and works regardless. But a Stratix-specific judge-snapshot comparison would be stronger because it captures the provider model version directly. The anchor set infers change through kappa on a proxy set.

Second: does the judge_snapshot.version field capture the Stratix judge version (incremented when a user edits a Stratix judge object) or the provider model version string (e.g. gpt-4o-2024-11-20)? The distinction determines whether the field solves the problem this piece describes or a different problem.

Third: is there a planned SDK endpoint to diff two trace-evaluation runs, analogous to comparisons.compare() for benchmarks? A native diff UX would eliminate the manual arithmetic a team currently performs.

These questions are open because the SDK exposure boundary and the version field semantics have not been confirmed against the shipping product. The anchor-set method works on any platform and answers the same question without waiting for them to close.Set Up Your Anchor Set This Afternoon

The method takes an afternoon. No new dependencies required.

Select 15 traces from this week's production traffic. Label them with pass or fail. Run your judge against them today and store the kappa value. Run the judge against them at the next deployment and compare kappa values. The first time kappa drops by a meaningful margin, you caught a judge-drift event that your dashboard would have shown as a score change with no explanation.

The cost of not detecting judge drift is false confidence in every score comparison. A team that deploys twice a day and compares eval scores between versions cannot distinguish improvement from drift without a fixed reference point. The anchor set is that reference point.

Register on Stratix to set up your first anchored eval comparison with versioned judges. The GEPA guide published at layerlens.ai covers recalibrating a drifted judge against current production distributions. The Education Portal has tutorials on configuring custom judges and interpreting judge_snapshot fields.