Earned Autonomy ¡ Part 2
The Agent That Did the Work Shouldn't Decide It's Done
Agents grade their own work generously. Why a separate judge with evidence and tools beats self-review, what it should see, and what it must never be told.
Josh McWilliam
10 min read
- agentic-systems
- harness
- claude-code
- production
- evals
Ask an agent whether its own work is finished and it will almost always say yes.
That isnât a character flaw. Itâs what the setup rewards. The agent that did the work is the one that wants to stop, it has read its own reasoning a hundred times, and itâs being asked to find the mistakes it didnât see the first time. If your agents mark their own work as done, you donât have a check. You have a formality.
The builder is the most generous grader in the room
Anthropicâs engineers put it plainly in Harness design for long-running application development (March 2026): âWhen asked to evaluate work theyâve produced, agents tend to respond by confidently praising the workâeven when, to a human observer, the quality is obviously mediocre.â
Their fix wasnât a better prompt. It was structure. âSeparating the agent doing the work from the agent judging it proves to be a strong lever,â and âtuning a standalone evaluator to be skeptical turns out to be far more tractable than making a generator critical of its own work, and once that external feedback exists, the generator has something concrete to iterate against.â
The Claude Code best practices say the same thing from the other side: âClaude stops when the work looks done. Without a check it can run, âlooks doneâ is the only signal available, and you become the verification loop.â And on review: âA fresh context improves code review since Claude wonât be biased toward code it just wrote.â
Iâve written before about giving the agent a check it can run. This post is about the next question: when the check is a judgment call, who makes it?
Four reasons self-review fails
The research behind this is consistent, and it points at four mechanisms that reinforce each other. Iâm quoting the papersâ own abstracts here, not my reading of their tables.
1. Finding a mistake is harder than fixing one. Huang and colleagues found that models âstruggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correctionâ (October 2023). Tyen and colleagues located the cause: poor self-correction âstems from LLMsâ inability to find logical mistakes, rather than their ability to correct a known mistakeâ (November 2023). For code, Olausson and colleagues found self-repair âbottlenecked by the modelâs ability to provide feedback on its own code,â with substantially larger gains when a stronger model wrote the feedback (June 2023). And Self-Correction Bench (July 2025) measured a 64.5% âSelf-Correction Blind Spotâ across 14 open-source non-reasoning models: they correct an error when itâs presented as someone elseâs and miss the identical error when itâs their own.
2. Judges favor what they recognize. Panickssery and colleagues found models have ânon-trivial accuracy at distinguishing themselves from other LLMs and humans,â and a âlinear correlation between self-recognition capability and the strength of self-preference biasâ (April 2024). The sharper finding comes from Chen and colleagues: harmful self-preference âpersists when evaluator models err as generators,â and âstronger models struggle more to recognize when they are wrongâ (April 2025).
Read those two together. Self-review is weakest on exactly the work where the model got it wrong, which is the only work a validator exists to catch.
3. Shared context lets the pair game the score. When the same model generates and rates, the evaluatorâs ratings can climb âwhile the generation quality remains stagnant or even decreases as judged by actual user preference.â Pan and colleagues found two factors that affect how bad that gets: model size, and âcontext sharing between the generator and the evaluatorâ (July 2024).
4. The agent that wants to stop is the one deciding. Anthropic watched their own evaluator do this: âI watched it identify legitimate issues, then talk itself into deciding they werenât a big deal and approve the work anyway.â
Separation is necessary, not sufficient
That last quote is about the separate evaluator, not the builder. Itâs the caveat that matters most.
âOut of the box, Claude is a poor QA agent,â the same post says. It âtended to test superficially, rather than probing edge cases.â The fix was a tuning loop: âread the evaluatorâs logs, find examples where its judgment diverged from mine,â and rewrite its prompt to close the gap. It took âseveral roundsâ before the grading looked reasonable.
What made that evaluator worth tuning was that it could check things itself. It used the Playwright MCP âto click through the running application the way a user would.â Its findings read like bug reports, not opinions: a route defined in the wrong order, so the framework parsed the word âreorderâ as an ID and returned an error.
The research points the same way. In the Agent-as-a-Judge work (October 2024), a judge with agentic tools, tested on a benchmark of 55 realistic AI development tasks, âdramatically outperforms LLM-as-a-Judge and is as reliable as our human evaluation baseline.â
So there are two different fixes for two different problems:
- A fresh context removes the shared-context problem. The judge hasnât read the builderâs reasoning, so it canât be talked into it.
- A different model, or a check that doesnât depend on a model at all, addresses familiarity. Verga and colleagues found a panel of smaller judges from âdisjoint model familiesâ outperformed a single large judge while being âover seven times less expensiveâ (April 2024).
A fresh context with the same model is a real improvement. It isnât the whole answer.
What the judge sees, and what it never sees
If you take one thing from this post, take this part. A separate validator is only as independent as its inputs.
It should get four things:
- The artifact. The diff, the changed files, and how to build or run it.
- The contract. The acceptance criteria, the out-of-scope list, the projectâs rules, and what counts as blocking.
- Evidence the harness captured itself. Raw test output, exit codes, screenshots, written by the system, not summarized by the agent.
- Tools to gather more. Read access to the repository, a test runner, a browser.
It should never get the claim.
Addy Osmaniâs doubt-driven development skill says it better than I can: âA fresh-context reviewer needs the artifact and the contract, not the journey.â And: âStrip your reasoning. If you hand over conclusions, youâll get back validation of your conclusions.â
The Claude Code guidance describes the same boundary: a reviewer in a fresh subagent âsees only the diff and the criteria you give it, not the reasoning that produced the change, so it evaluates the result on its own terms.â
Thereâs one nuance. Anthropicâs own code-review plugin tells each reviewer the pull requestâs title and description, to âprovide context regarding the authorâs intent.â I donât think that contradicts the rule. Intent is part of the contract. âI verified the edge casesâ and âall tests passâ are claims. The line I use:
Give the judge what should be true. Withhold why the author believes itâs true.
Evidence has to be produced, not reported
Claude Codeâs /goal command is a useful illustration of the limit. Its evaluator judges the condition against what Claude has surfaced in the conversation, and the docs are upfront that âit doesnât run commands or read files independently.â Thatâs fine for what /goal is for. But a judge that only reads the transcript can only be as good as what the worker chose to show it.
A gating validator should either run the checks itself or read evidence files the harness wrote. Anthropicâs Code Review documentation suggests a rule I like: âbehavior claims need a file:line citation in the source, not an inference from naming.â
On this site, most of the judging isnât done by a model at all. One command runs the type check, the full build, and a content guard that fails the build if any page contains a claim Iâve decided not to make. None of those can be argued with. The model-based judgment sits on top of checks like that, not instead of them.
The verdict needs an âI donât knowâ
A judge forced to choose between pass and fail will sometimes guess. Anthropicâs guide to agent evals (January 2026) recommends giving it a way out, âlike providing an instruction to return âUnknownâ when it doesnât have enough information.â
Claude Codeâs workflows handle the same problem in research: when verifier agents canât check a claim, âthe report lists that claim as unverified instead of counting it as refuted.â For a gate, Iâd go one step further. Unknown and unverified are their own states, and neither one counts as a pass.
Tell it what not to flag
A judge can fail in the other direction too. The Claude Code guidance warns that âa reviewer prompted to find gaps will usually report some, even when the work is sound, because that is what it was asked to do,â and that chasing every finding âleads to over-engineering.â Its advice: âflag only gaps that affect correctness or the stated requirements, and treat the rest as optional.â
Thatâs the scope I want from a validator. It should be strict about the contract and quiet about taste.
Two or three rounds, then stop
On how many fix-and-recheck cycles to allow, the practitioners I trust land in the same place.
- Osmaniâs skill stops at three cycles (âescalate to user, donât grind a fourth aloneâ), and if three feels obviously too few, âthe artifact is too big â return to Step 2 and decompose. Do not lift the bound.â It also names âdoubt theaterâ: two or more cycles in which the reviewer raised substantive findings and none were treated as actionable, a sign youâre âvalidating, not doubting.â
- The Claude Code guidance: âAfter two failed corrections, /clear and write a better initial prompt incorporating what you learned.â
- The Code Review docs suggest a rule that, after the first review, suppresses new nits and posts only important findings, which âstops a one-line fix from reaching round seven on style alone.â
One more detail from Anthropicâs harness post changed how I think about loops. On subjective work, scores generally improved over iterations, but âI regularly saw cases where I preferred a middle iteration over the last one.â So keep the best checkpoint, not the most recent one.
What it costs
Less than youâd think, relative to what it protects.
In Anthropicâs simplified harness run, three QA rounds cost $10.39 of a $124.70 run, about 8%, and about 25 of roughly 230 minutes. The three build rounds were most of the spend. Thatâs the point. And the QA caught what the builder had left undone: âAudio recording is still stub-only,â and core features that were âdisplay-only without interactive depth.â
The post is careful about when this pays: the evaluator âis worth the cost when the task sits beyond what the current model does reliably solo.â On easy work, itâs overhead.
Anthropicâs managed Code Review (March 2026) is the other data point Iâve seen. Reviews average $15â25. Internally, the share of pull requests getting substantive review comments went from 16% to 54%, and âless than 1% of findings are marked incorrect.â Thatâs vendor-reported, and it measures precision: how often the reviewer is wrong when it speaks. It doesnât tell you what it missed. For a gate, what it misses is the number that matters.
What this changes in Faber
Faber runs work through Frame, Architect, Build, Evaluate, and Release, with validators between phases. Some of its packs already run their validators as separate agents. The work now is making that a guarantee every pack gets, including the software pack, and making sure a validatorâs verdict actually decides what happens next. The hardening spec lays it out; hereâs where each piece stands.
- Shipped: the CLI runtime runs the loop as plain code and starts a fresh agent session for every step, so a validator step there already gets its own context (PR #220).
- Merged, in the next release: verdicts that gate the run. Each step returns a structured verdict the runtime parses, and a failure stops the run instead of being logged and passed over (PR #243).
- Planned: a validator runner with the four-part input packet above. Stated intent is allowed; the makerâs reasoning and claims are withheld. Validators get read-only tools plus the ability to run checks, and read evidence files the harness writes.
- Planned: per-criterion verdicts of pass, fail, or unknown, with unknown never counting as a pass, and findings limited to correctness and the stated criteria.
- Planned: a bounded fix loop of at most three cycles, a fresh maker session after two failed fixes, a stop when nothing is changing, and the best checkpoint kept rather than the last.
- Planned: model and provider set per validator, so a high-risk change, like authentication, a data migration, or a release step, can be judged by a different model family than the one that wrote it.
None of this replaces my review. I still read every diff before it merges, as I wrote in Building the Fractary Platform. What changes is what reaches me: work that something other than its author has already tried to break.
This is the second post in the Earned Autonomy series. Previous: Who Decides an Agentâs Work Is Done? How Faber Is Changing. Next: Done Is a Contract: Four Layers Your Agents Canât Rewrite.