Earned Autonomy ¡ Part 6
A Validator Is Not an Eval: How to Measure the Judge
A validator decides whether one agent run passes. It can't tell you whether the pipeline got better. How to build evals that measure the judge, not the work.
Josh McWilliam
9 min read
- evals
- harness
- agentic-systems
- production
A passing validator tells you almost nothing about whether your agent pipeline is any good.
It tells you that one run cleared one gate. Whether the gate itself is any good, whether it stops bad work or waves it through, is a different question, and the validator canât answer it about itself.
I spent the earlier posts in this series arguing for validators: separate judges, written contracts, evidence instead of claims. This post is about the thing that keeps all of that honest. You have to measure the judge.
Two jobs that look alike
Hamel Husain and Shreya Shankar draw the line cleanly in their evals FAQ (June 2025). âGuardrails are inline safety checks that sit directly in the request/response path,â and with a guardrail, âfalse positives are treated as production bugs.â
Evaluators are the other thing. They âtypically run after a response is produced,â their verdicts feed âdashboards, regression tests, and model-improvement loops,â and they âdo not block the original answer.â
The validators in an agent pipeline are a hybrid. Theyâre LLM judges, which makes them look like evaluators, but they sit in the path and block, which makes them guardrails.
That hybrid has a consequence people skip. A blocking judge is a component with an error rate. It isnât a measurement of the system. Itâs part of the system, and it needs measuring like any other part.
Why the gate canât grade itself
A validator judges one run. It canât tell you whether your pipeline improved after you changed a prompt, upgraded a model, or rewrote the validator. Those are questions about many runs, compared before and after.
The tempting shortcut is to reuse the validator as the grader for your eval suite. It already knows the criteria. It already runs.
Donât. Anthropicâs guide to agent evals (January 2026) borrows a picture from safety engineering: âLike the Swiss Cheese Model from safety engineering, no single evaluation layer catches every issue.â Layers work because their holes are in different places.
Grade your evals with your own gate and the holes line up. A lenient validator passes bad work in the run, then scores that same run as a success in the eval. The gate and the score both look fine, and both are wrong in the same direction.
An eval needs ground truth the pipeline didnât produce: hidden tests the agent never saw, reference solutions, and labels from a person.
Start with the failures you already have
You donât need a benchmark to start. You need your own failures.
Husainâs error-analysis guide starts with about a hundred representative traces, read by a person, with notes grouped into failure categories. It stops when â~100 traces reviewed with no new failure types appearing in the last 20.â Thatâs the point where reading more stops teaching you anything.
Then turn the failures into tasks. Anthropicâs guide is direct about size: âIn reality, 20-50 simple tasks drawn from real failures is a great start.â Early on, changes have a big effect, and âthis large effect size means small sample sizes suffice.â
Each task needs two things. First, âa reference solution: a known working output that passes all graders.â If no correct answer passes your graders, the task is broken, not the agent.
Second, a task has to be unambiguous: âA good task is one where two domain experts would independently reach the same pass/fail verdict.â If two people would argue about whether the output passed, the agent canât be graded on it either.
I already have a version of this list. The content guard on this site fails the build on strings Iâve decided never to ship, and as I wrote in Vibe Coding Gets You a Demo, every rule in it is a mistake that already happened once. Thatâs what a regression suite is: past failures, written down so they canât come back quietly.
Run it like you mean it
How you run the suite matters as much as whatâs in it. Three rules from the Anthropic guide changed how I think about it.
Start every trial clean. Anthropic âobserved Claude gaining an unfair advantage on some tasks by examining the git history from previous trials.â The fix: âEach trial should be âisolatedâ by starting from a clean environment.â For a coding pipeline, that means a fresh checkout per trial, with nothing left over from the last attempt.
Grade the outcome, not the route. Agents find valid approaches the evalâs author didnât anticipate. âSo as not to unnecessarily punish creativity, itâs often better to grade what the agent produced, not the path it took.â Check that the feature works, not that the agent called the tools in the order you would have.
Keep two suites. Capability evals ask what the agent can do, and âshould start at a low pass rate,â which gives you something to climb. Regression evals ask whether it still does what it used to, and âshould have a nearly 100% pass rate.â A capability task the agent passes every time graduates into the regression suite.
And run each task more than once. Agents arenât deterministic, and one pass tells you little. How many runs, and what to count, is the subject of the next post.
Read the transcripts
The rule from the Anthropic guide Iâd keep if I could keep only one: âAs a rule, we do not take eval scores at face value until someone digs into the details of the eval and reads some transcripts.â
Graders fail too, and the failures are not subtle once you look. In the same guide, a frontier model âinitially scored 42% on CORE-Benchâ until a researcher found the problems: rigid grading that penalized â96.12â when expecting â96.124991âŚâ, ambiguous task specs, and stochastic tasks that were impossible to reproduce exactly. After the fixes and a less constrained harness, the score was 95%.
The model didnât change. The measurement did.
OpenAI found the same thing at a larger scale. In February 2026 it stopped reporting SWE-bench Verified, a benchmark it had created. It audited 138 problems that a model didnât consistently solve, and â59.4% of the 138 problems contained material issues in test design and/or problem description.â It also found signs that frontier models had seen some of the problems in training.
A score is a claim about two things at once: the agent and the grader. Reading transcripts is how you find out which one youâre looking at.
Measure the judge
If the validator is a component, it gets tested like one. There are two ways to do it, and Iâd do both.
Against human labels. Husainâs guide to validating an evaluator is the most concrete protocol Iâve found. Label a set of outputs pass or fail yourself. Split them into training (10â20%, used as examples in the judgeâs prompt), dev (40â45%, used to tune it) and test (40â45%, held back).
Tune the judge on the dev set until both of its rates clear the target: âTPR > 90% AND TNR > 90%.â The true positive rate is how often the judge says pass when you said pass. The true negative rate is how often it says fail when you said fail.
Then run it on the test set exactly once. âDo not iterate after seeing test set results.â And âpin exact model versionsâ for the judge, because âproviders update models without notice, causing silent drift.â
For a gate, the true negative rate is the one that protects you. Itâs the share of bad work the validator actually stops. A judge with a great pass rate on good work and a weak catch rate on bad work feels smooth and ships defects.
Against planted defects. OpenAIâs CriticGPT work (June 2024) trained a model to critique code, and to train and test it, âwe asked AI trainers to manually insert these mistakes into code written by ChatGPT.â Then they checked whether the critic caught the inserted bugs.
The same trick works on an agent pipeline, and itâs cheaper than labeling. Take a change you know is correct and break it on purpose: drop one of the acceptance criteria, weaken a test so it canât fail, add something that was explicitly out of scope. Then see whether the validator fails it.
That catches the failure that matters most. Qodoâs code review benchmark, which it built âby injecting verified bugs and best practice violations into real, merged pull requests,â reports that âmost tools cluster toward high precision and low recall.â Qodo runs the benchmark and ranks its own product, so read it as a vendorâs result. The pattern it describes is still the one to test for: a reviewer thatâs rarely wrong about what it flags and quietly misses most of what it doesnât.
Precision on its own can look excellent. Anthropic says of its managed Code Review (March 2026) that âless than 1% of findings are marked incorrect,â and that âbefore, 16% of PRs got substantive review comments. Now 54% do.â Those are good numbers about the findings it made. They donât say how many issues went unflagged, and for a gate, thatâs the number you need next to it.
Does each piece earn its keep?
Measuring the judge tells you whether itâs accurate. It doesnât tell you whether itâs worth having.
Claude Codeâs plugin evals answer that with an ablation. âEach case runs three times by default,â and then the runs are repeated with no plugin loaded, so âyou get two scores, WITH and W/OUT. Their difference, Î, is what the plugin contributed. If a case scores 1.0 both with and without the plugin, the plugin isnât what made it pass.â
The same docs say to pin the model âso a model rollout isnât mistaken for a plugin regression.â Change one thing at a time, and know which thing you changed.
That design is the right test for every piece of an agent harness: each validator, each skill, each planning step. Run the suite with it and without it. If the scores match, the piece isnât helping, whatever it costs. And rerun the comparison at every model upgrade, because, as I argued in the previous post, a component that helped last quarterâs model may be dead weight for the next one.
What this changes in Faber
In Faber, validators are the gates between phases, and making every verdict actually stop a run is merged and in the next release. What Iâm adding is the measurement around them. Itâs all in the hardening spec, and itâs planned, not shipped:
- A pilot first. Ten past work items, three runs each, on the CLI runtime, which already starts every step in a fresh agent session. It measures pass rates, cost per verified success and how much the results vary, and those numbers set the size of the full suite and how much extra cost a change is allowed to add.
- Ground truth from outside the pipeline. An eval task is a frozen past work item: its requirements, its starting state, and a hidden answer key built from what a person accepted. LLM graders get calibrated against human labels. Nothing is graded by the pipelineâs own verdicts.
- Every validator measured before itâs trusted to gate. A catch rate on planted defects and a false-alarm rate on accepted outputs, with targets of at least 90% caught and at most 10% false alarms.
- Changes ship with evidence. A skill, prompt or model change comes with an eval comparison, and rejected changes get logged. Every scaffold is re-tested at every model upgrade, and the with-and-without comparison above is how I intend to do it.
When the pilot has run, Iâll publish what it showed: one practitionerâs numbers on my own work, including the parts that didnât hold up.
This is the sixth post in the Earned Autonomy series. Previous: Your Agent Harness Is a List of What the Model Canât Do. Next: pass^k: The Reliability Math of Unattended Agents.