← Blog

Earned Autonomy ¡ Part 6

A Validator Is Not an Eval: How to Measure the Judge

A validator decides whether one agent run passes. It can't tell you whether the pipeline got better. How to build evals that measure the judge, not the work.

Josh McWilliam

9 min read

  • evals
  • harness
  • agentic-systems
  • production

A passing validator tells you almost nothing about whether your agent pipeline is any good.

It tells you that one run cleared one gate. Whether the gate itself is any good, whether it stops bad work or waves it through, is a different question, and the validator can’t answer it about itself.

I spent the earlier posts in this series arguing for validators: separate judges, written contracts, evidence instead of claims. This post is about the thing that keeps all of that honest. You have to measure the judge.

Two jobs that look alike

Hamel Husain and Shreya Shankar draw the line cleanly in their evals FAQ (June 2025). “Guardrails are inline safety checks that sit directly in the request/response path,” and with a guardrail, “false positives are treated as production bugs.”

Evaluators are the other thing. They “typically run after a response is produced,” their verdicts feed “dashboards, regression tests, and model-improvement loops,” and they “do not block the original answer.”

The validators in an agent pipeline are a hybrid. They’re LLM judges, which makes them look like evaluators, but they sit in the path and block, which makes them guardrails.

That hybrid has a consequence people skip. A blocking judge is a component with an error rate. It isn’t a measurement of the system. It’s part of the system, and it needs measuring like any other part.

Why the gate can’t grade itself

A validator judges one run. It can’t tell you whether your pipeline improved after you changed a prompt, upgraded a model, or rewrote the validator. Those are questions about many runs, compared before and after.

The tempting shortcut is to reuse the validator as the grader for your eval suite. It already knows the criteria. It already runs.

Don’t. Anthropic’s guide to agent evals (January 2026) borrows a picture from safety engineering: “Like the Swiss Cheese Model from safety engineering, no single evaluation layer catches every issue.” Layers work because their holes are in different places.

Grade your evals with your own gate and the holes line up. A lenient validator passes bad work in the run, then scores that same run as a success in the eval. The gate and the score both look fine, and both are wrong in the same direction.

An eval needs ground truth the pipeline didn’t produce: hidden tests the agent never saw, reference solutions, and labels from a person.

Start with the failures you already have

You don’t need a benchmark to start. You need your own failures.

Husain’s error-analysis guide starts with about a hundred representative traces, read by a person, with notes grouped into failure categories. It stops when “~100 traces reviewed with no new failure types appearing in the last 20.” That’s the point where reading more stops teaching you anything.

Then turn the failures into tasks. Anthropic’s guide is direct about size: “In reality, 20-50 simple tasks drawn from real failures is a great start.” Early on, changes have a big effect, and “this large effect size means small sample sizes suffice.”

Each task needs two things. First, “a reference solution: a known working output that passes all graders.” If no correct answer passes your graders, the task is broken, not the agent.

Second, a task has to be unambiguous: “A good task is one where two domain experts would independently reach the same pass/fail verdict.” If two people would argue about whether the output passed, the agent can’t be graded on it either.

I already have a version of this list. The content guard on this site fails the build on strings I’ve decided never to ship, and as I wrote in Vibe Coding Gets You a Demo, every rule in it is a mistake that already happened once. That’s what a regression suite is: past failures, written down so they can’t come back quietly.

Run it like you mean it

How you run the suite matters as much as what’s in it. Three rules from the Anthropic guide changed how I think about it.

Start every trial clean. Anthropic “observed Claude gaining an unfair advantage on some tasks by examining the git history from previous trials.” The fix: “Each trial should be ‘isolated’ by starting from a clean environment.” For a coding pipeline, that means a fresh checkout per trial, with nothing left over from the last attempt.

Grade the outcome, not the route. Agents find valid approaches the eval’s author didn’t anticipate. “So as not to unnecessarily punish creativity, it’s often better to grade what the agent produced, not the path it took.” Check that the feature works, not that the agent called the tools in the order you would have.

Keep two suites. Capability evals ask what the agent can do, and “should start at a low pass rate,” which gives you something to climb. Regression evals ask whether it still does what it used to, and “should have a nearly 100% pass rate.” A capability task the agent passes every time graduates into the regression suite.

And run each task more than once. Agents aren’t deterministic, and one pass tells you little. How many runs, and what to count, is the subject of the next post.

Read the transcripts

The rule from the Anthropic guide I’d keep if I could keep only one: “As a rule, we do not take eval scores at face value until someone digs into the details of the eval and reads some transcripts.”

Graders fail too, and the failures are not subtle once you look. In the same guide, a frontier model “initially scored 42% on CORE-Bench” until a researcher found the problems: rigid grading that penalized “96.12” when expecting “96.124991…”, ambiguous task specs, and stochastic tasks that were impossible to reproduce exactly. After the fixes and a less constrained harness, the score was 95%.

The model didn’t change. The measurement did.

OpenAI found the same thing at a larger scale. In February 2026 it stopped reporting SWE-bench Verified, a benchmark it had created. It audited 138 problems that a model didn’t consistently solve, and “59.4% of the 138 problems contained material issues in test design and/or problem description.” It also found signs that frontier models had seen some of the problems in training.

A score is a claim about two things at once: the agent and the grader. Reading transcripts is how you find out which one you’re looking at.

Measure the judge

If the validator is a component, it gets tested like one. There are two ways to do it, and I’d do both.

Against human labels. Husain’s guide to validating an evaluator is the most concrete protocol I’ve found. Label a set of outputs pass or fail yourself. Split them into training (10–20%, used as examples in the judge’s prompt), dev (40–45%, used to tune it) and test (40–45%, held back).

Tune the judge on the dev set until both of its rates clear the target: “TPR > 90% AND TNR > 90%.” The true positive rate is how often the judge says pass when you said pass. The true negative rate is how often it says fail when you said fail.

Then run it on the test set exactly once. “Do not iterate after seeing test set results.” And “pin exact model versions” for the judge, because “providers update models without notice, causing silent drift.”

For a gate, the true negative rate is the one that protects you. It’s the share of bad work the validator actually stops. A judge with a great pass rate on good work and a weak catch rate on bad work feels smooth and ships defects.

Against planted defects. OpenAI’s CriticGPT work (June 2024) trained a model to critique code, and to train and test it, “we asked AI trainers to manually insert these mistakes into code written by ChatGPT.” Then they checked whether the critic caught the inserted bugs.

The same trick works on an agent pipeline, and it’s cheaper than labeling. Take a change you know is correct and break it on purpose: drop one of the acceptance criteria, weaken a test so it can’t fail, add something that was explicitly out of scope. Then see whether the validator fails it.

That catches the failure that matters most. Qodo’s code review benchmark, which it built “by injecting verified bugs and best practice violations into real, merged pull requests,” reports that “most tools cluster toward high precision and low recall.” Qodo runs the benchmark and ranks its own product, so read it as a vendor’s result. The pattern it describes is still the one to test for: a reviewer that’s rarely wrong about what it flags and quietly misses most of what it doesn’t.

Precision on its own can look excellent. Anthropic says of its managed Code Review (March 2026) that “less than 1% of findings are marked incorrect,” and that “before, 16% of PRs got substantive review comments. Now 54% do.” Those are good numbers about the findings it made. They don’t say how many issues went unflagged, and for a gate, that’s the number you need next to it.

Does each piece earn its keep?

Measuring the judge tells you whether it’s accurate. It doesn’t tell you whether it’s worth having.

Claude Code’s plugin evals answer that with an ablation. “Each case runs three times by default,” and then the runs are repeated with no plugin loaded, so “you get two scores, WITH and W/OUT. Their difference, Δ, is what the plugin contributed. If a case scores 1.0 both with and without the plugin, the plugin isn’t what made it pass.”

The same docs say to pin the model “so a model rollout isn’t mistaken for a plugin regression.” Change one thing at a time, and know which thing you changed.

That design is the right test for every piece of an agent harness: each validator, each skill, each planning step. Run the suite with it and without it. If the scores match, the piece isn’t helping, whatever it costs. And rerun the comparison at every model upgrade, because, as I argued in the previous post, a component that helped last quarter’s model may be dead weight for the next one.

What this changes in Faber

In Faber, validators are the gates between phases, and making every verdict actually stop a run is merged and in the next release. What I’m adding is the measurement around them. It’s all in the hardening spec, and it’s planned, not shipped:

  • A pilot first. Ten past work items, three runs each, on the CLI runtime, which already starts every step in a fresh agent session. It measures pass rates, cost per verified success and how much the results vary, and those numbers set the size of the full suite and how much extra cost a change is allowed to add.
  • Ground truth from outside the pipeline. An eval task is a frozen past work item: its requirements, its starting state, and a hidden answer key built from what a person accepted. LLM graders get calibrated against human labels. Nothing is graded by the pipeline’s own verdicts.
  • Every validator measured before it’s trusted to gate. A catch rate on planted defects and a false-alarm rate on accepted outputs, with targets of at least 90% caught and at most 10% false alarms.
  • Changes ship with evidence. A skill, prompt or model change comes with an eval comparison, and rejected changes get logged. Every scaffold is re-tested at every model upgrade, and the with-and-without comparison above is how I intend to do it.

When the pilot has run, I’ll publish what it showed: one practitioner’s numbers on my own work, including the parts that didn’t hold up.


This is the sixth post in the Earned Autonomy series. Previous: Your Agent Harness Is a List of What the Model Can’t Do. Next: pass^k: The Reliability Math of Unattended Agents.

The Stack

Build on the stack the studio builds with.

Open-source, Apache-2.0 tools for building agentic systems without vendor lock-in. The same stack behind every venture on this site.