← Blog

Earned Autonomy ¡ Part 2

The Agent That Did the Work Shouldn't Decide It's Done

Agents grade their own work generously. Why a separate judge with evidence and tools beats self-review, what it should see, and what it must never be told.

Josh McWilliam

10 min read

  • agentic-systems
  • harness
  • claude-code
  • production
  • evals

Ask an agent whether its own work is finished and it will almost always say yes.

That isn’t a character flaw. It’s what the setup rewards. The agent that did the work is the one that wants to stop, it has read its own reasoning a hundred times, and it’s being asked to find the mistakes it didn’t see the first time. If your agents mark their own work as done, you don’t have a check. You have a formality.

The builder is the most generous grader in the room

Anthropic’s engineers put it plainly in Harness design for long-running application development (March 2026): “When asked to evaluate work they’ve produced, agents tend to respond by confidently praising the work—even when, to a human observer, the quality is obviously mediocre.”

Their fix wasn’t a better prompt. It was structure. “Separating the agent doing the work from the agent judging it proves to be a strong lever,” and “tuning a standalone evaluator to be skeptical turns out to be far more tractable than making a generator critical of its own work, and once that external feedback exists, the generator has something concrete to iterate against.”

The Claude Code best practices say the same thing from the other side: “Claude stops when the work looks done. Without a check it can run, ‘looks done’ is the only signal available, and you become the verification loop.” And on review: “A fresh context improves code review since Claude won’t be biased toward code it just wrote.”

I’ve written before about giving the agent a check it can run. This post is about the next question: when the check is a judgment call, who makes it?

Four reasons self-review fails

The research behind this is consistent, and it points at four mechanisms that reinforce each other. I’m quoting the papers’ own abstracts here, not my reading of their tables.

1. Finding a mistake is harder than fixing one. Huang and colleagues found that models “struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction” (October 2023). Tyen and colleagues located the cause: poor self-correction “stems from LLMs’ inability to find logical mistakes, rather than their ability to correct a known mistake” (November 2023). For code, Olausson and colleagues found self-repair “bottlenecked by the model’s ability to provide feedback on its own code,” with substantially larger gains when a stronger model wrote the feedback (June 2023). And Self-Correction Bench (July 2025) measured a 64.5% “Self-Correction Blind Spot” across 14 open-source non-reasoning models: they correct an error when it’s presented as someone else’s and miss the identical error when it’s their own.

2. Judges favor what they recognize. Panickssery and colleagues found models have “non-trivial accuracy at distinguishing themselves from other LLMs and humans,” and a “linear correlation between self-recognition capability and the strength of self-preference bias” (April 2024). The sharper finding comes from Chen and colleagues: harmful self-preference “persists when evaluator models err as generators,” and “stronger models struggle more to recognize when they are wrong” (April 2025).

Read those two together. Self-review is weakest on exactly the work where the model got it wrong, which is the only work a validator exists to catch.

3. Shared context lets the pair game the score. When the same model generates and rates, the evaluator’s ratings can climb “while the generation quality remains stagnant or even decreases as judged by actual user preference.” Pan and colleagues found two factors that affect how bad that gets: model size, and “context sharing between the generator and the evaluator” (July 2024).

4. The agent that wants to stop is the one deciding. Anthropic watched their own evaluator do this: “I watched it identify legitimate issues, then talk itself into deciding they weren’t a big deal and approve the work anyway.”

Separation is necessary, not sufficient

That last quote is about the separate evaluator, not the builder. It’s the caveat that matters most.

“Out of the box, Claude is a poor QA agent,” the same post says. It “tended to test superficially, rather than probing edge cases.” The fix was a tuning loop: “read the evaluator’s logs, find examples where its judgment diverged from mine,” and rewrite its prompt to close the gap. It took “several rounds” before the grading looked reasonable.

What made that evaluator worth tuning was that it could check things itself. It used the Playwright MCP “to click through the running application the way a user would.” Its findings read like bug reports, not opinions: a route defined in the wrong order, so the framework parsed the word “reorder” as an ID and returned an error.

The research points the same way. In the Agent-as-a-Judge work (October 2024), a judge with agentic tools, tested on a benchmark of 55 realistic AI development tasks, “dramatically outperforms LLM-as-a-Judge and is as reliable as our human evaluation baseline.”

So there are two different fixes for two different problems:

  • A fresh context removes the shared-context problem. The judge hasn’t read the builder’s reasoning, so it can’t be talked into it.
  • A different model, or a check that doesn’t depend on a model at all, addresses familiarity. Verga and colleagues found a panel of smaller judges from “disjoint model families” outperformed a single large judge while being “over seven times less expensive” (April 2024).

A fresh context with the same model is a real improvement. It isn’t the whole answer.

What the judge sees, and what it never sees

If you take one thing from this post, take this part. A separate validator is only as independent as its inputs.

It should get four things:

  1. The artifact. The diff, the changed files, and how to build or run it.
  2. The contract. The acceptance criteria, the out-of-scope list, the project’s rules, and what counts as blocking.
  3. Evidence the harness captured itself. Raw test output, exit codes, screenshots, written by the system, not summarized by the agent.
  4. Tools to gather more. Read access to the repository, a test runner, a browser.

It should never get the claim.

Addy Osmani’s doubt-driven development skill says it better than I can: “A fresh-context reviewer needs the artifact and the contract, not the journey.” And: “Strip your reasoning. If you hand over conclusions, you’ll get back validation of your conclusions.”

The Claude Code guidance describes the same boundary: a reviewer in a fresh subagent “sees only the diff and the criteria you give it, not the reasoning that produced the change, so it evaluates the result on its own terms.”

There’s one nuance. Anthropic’s own code-review plugin tells each reviewer the pull request’s title and description, to “provide context regarding the author’s intent.” I don’t think that contradicts the rule. Intent is part of the contract. “I verified the edge cases” and “all tests pass” are claims. The line I use:

Give the judge what should be true. Withhold why the author believes it’s true.

Evidence has to be produced, not reported

Claude Code’s /goal command is a useful illustration of the limit. Its evaluator judges the condition against what Claude has surfaced in the conversation, and the docs are upfront that “it doesn’t run commands or read files independently.” That’s fine for what /goal is for. But a judge that only reads the transcript can only be as good as what the worker chose to show it.

A gating validator should either run the checks itself or read evidence files the harness wrote. Anthropic’s Code Review documentation suggests a rule I like: “behavior claims need a file:line citation in the source, not an inference from naming.”

On this site, most of the judging isn’t done by a model at all. One command runs the type check, the full build, and a content guard that fails the build if any page contains a claim I’ve decided not to make. None of those can be argued with. The model-based judgment sits on top of checks like that, not instead of them.

The verdict needs an “I don’t know”

A judge forced to choose between pass and fail will sometimes guess. Anthropic’s guide to agent evals (January 2026) recommends giving it a way out, “like providing an instruction to return ‘Unknown’ when it doesn’t have enough information.”

Claude Code’s workflows handle the same problem in research: when verifier agents can’t check a claim, “the report lists that claim as unverified instead of counting it as refuted.” For a gate, I’d go one step further. Unknown and unverified are their own states, and neither one counts as a pass.

Tell it what not to flag

A judge can fail in the other direction too. The Claude Code guidance warns that “a reviewer prompted to find gaps will usually report some, even when the work is sound, because that is what it was asked to do,” and that chasing every finding “leads to over-engineering.” Its advice: “flag only gaps that affect correctness or the stated requirements, and treat the rest as optional.”

That’s the scope I want from a validator. It should be strict about the contract and quiet about taste.

Two or three rounds, then stop

On how many fix-and-recheck cycles to allow, the practitioners I trust land in the same place.

  • Osmani’s skill stops at three cycles (“escalate to user, don’t grind a fourth alone”), and if three feels obviously too few, “the artifact is too big — return to Step 2 and decompose. Do not lift the bound.” It also names “doubt theater”: two or more cycles in which the reviewer raised substantive findings and none were treated as actionable, a sign you’re “validating, not doubting.”
  • The Claude Code guidance: “After two failed corrections, /clear and write a better initial prompt incorporating what you learned.”
  • The Code Review docs suggest a rule that, after the first review, suppresses new nits and posts only important findings, which “stops a one-line fix from reaching round seven on style alone.”

One more detail from Anthropic’s harness post changed how I think about loops. On subjective work, scores generally improved over iterations, but “I regularly saw cases where I preferred a middle iteration over the last one.” So keep the best checkpoint, not the most recent one.

What it costs

Less than you’d think, relative to what it protects.

In Anthropic’s simplified harness run, three QA rounds cost $10.39 of a $124.70 run, about 8%, and about 25 of roughly 230 minutes. The three build rounds were most of the spend. That’s the point. And the QA caught what the builder had left undone: “Audio recording is still stub-only,” and core features that were “display-only without interactive depth.”

The post is careful about when this pays: the evaluator “is worth the cost when the task sits beyond what the current model does reliably solo.” On easy work, it’s overhead.

Anthropic’s managed Code Review (March 2026) is the other data point I’ve seen. Reviews average $15–25. Internally, the share of pull requests getting substantive review comments went from 16% to 54%, and “less than 1% of findings are marked incorrect.” That’s vendor-reported, and it measures precision: how often the reviewer is wrong when it speaks. It doesn’t tell you what it missed. For a gate, what it misses is the number that matters.

What this changes in Faber

Faber runs work through Frame, Architect, Build, Evaluate, and Release, with validators between phases. Some of its packs already run their validators as separate agents. The work now is making that a guarantee every pack gets, including the software pack, and making sure a validator’s verdict actually decides what happens next. The hardening spec lays it out; here’s where each piece stands.

  • Shipped: the CLI runtime runs the loop as plain code and starts a fresh agent session for every step, so a validator step there already gets its own context (PR #220).
  • Merged, in the next release: verdicts that gate the run. Each step returns a structured verdict the runtime parses, and a failure stops the run instead of being logged and passed over (PR #243).
  • Planned: a validator runner with the four-part input packet above. Stated intent is allowed; the maker’s reasoning and claims are withheld. Validators get read-only tools plus the ability to run checks, and read evidence files the harness writes.
  • Planned: per-criterion verdicts of pass, fail, or unknown, with unknown never counting as a pass, and findings limited to correctness and the stated criteria.
  • Planned: a bounded fix loop of at most three cycles, a fresh maker session after two failed fixes, a stop when nothing is changing, and the best checkpoint kept rather than the last.
  • Planned: model and provider set per validator, so a high-risk change, like authentication, a data migration, or a release step, can be judged by a different model family than the one that wrote it.

None of this replaces my review. I still read every diff before it merges, as I wrote in Building the Fractary Platform. What changes is what reaches me: work that something other than its author has already tried to break.


This is the second post in the Earned Autonomy series. Previous: Who Decides an Agent’s Work Is Done? How Faber Is Changing. Next: Done Is a Contract: Four Layers Your Agents Can’t Rewrite.

The Stack

Build on the stack the studio builds with.

Open-source, Apache-2.0 tools for building agentic systems without vendor lock-in. The same stack behind every venture on this site.