Earned Autonomy ¡ Part 1
Who Decides an Agent's Work Is Done? How Faber Is Changing
Agents got good at producing work. Deciding it's done is the hard part now. What the field learned about loops, judges, and evals, and how I'm changing Faber.
Josh McWilliam
9 min read
- agentic-systems
- harness
- evals
- platform
- open-source
Getting an agent to produce work is no longer the hard part. Knowing when to believe it is.
Ask an agent whether it finished the job and it will almost always say yes. Anthropicâs engineers put it bluntly in a post on harness design (March 2026): when asked to evaluate work theyâve produced, agents âtend to respond by confidently praising the workâeven when, to a human observer, the quality is obviously mediocre.â
So the question I care about most has changed. It used to be âcan the agent build this?â Now itâs âwho decides itâs done, and on what evidence?â
Why I went looking
Every change in the studio runs through five phases: Frame, Architect, Build, Evaluate, Release. Thatâs the method the studioâs Faber tooling encodes, and I described it in Vibe Coding Gets You a Demo with one sentence I still stand behind: âThe agent that builds a thing is not the one that decides itâs done.â
Thatâs a principle. I wanted to know what it takes in practice.
So I sat down to revamp Faber and its software pack, faber-code, against what the people building these systems have published: the labsâ engineering write-ups, the documentation for the agent tools I use daily, the open-source spec-driven tools, and the research on why models are bad at grading their own work. The result was a long research note and a hardening spec for Faber.
This series is the readable version. The full research note has every source, and the hardening spec is public on GitHub.
Where Faber started
Faber wasnât starting from zero. The phases were there. Approvals at the boundaries were there. So was the path from an issue to a pull request, with workflows that inherit from a shared core so each pack only adds whatâs specific to its domain.
One important piece had already shipped: a runtime that runs the loop as plain code. fractary-faber workflow-execute walks the plan step by step and starts a fresh agent session for each one (PR #220). No master agent decides what happens next. Code does.
What the research changed is narrower, and it matters more. In most agent tooling, mine included, âdoneâ has meant the steps ran. The goal is for âdoneâ to mean thereâs evidence the result works, and that something other than the maker checked it.
Anthropicâs Claude Code documentation describes the default failure in two sentences in its best practices: âClaude stops when the work looks done. Without a check it can run, âlooks doneâ is the only signal available.â
Everything below is about replacing âlooks doneâ with something better.
Four shifts the field converged on
I read a lot of sources that disagree about a lot of things. On these four, they line up.
1. Code owns the loop. The model judges inside each step.
The control-flow question is mostly settled. Deterministic code holds the outer loop: state, stage transitions, budgets, gates, and human sign-off. The model does the judgment inside each step.
Claude Codeâs workflow documentation states it plainly: âA workflow script holds the loop, the branching, and the intermediate results itself, so Claudeâs context holds only the final answer.â The runtime even makes timestamps and random numbers throw inside a script âso that a relaunched run repeats the same agent() calls.â
The twist is who writes the script. Increasingly itâs the model. The model writes the graph; a deterministic runtime executes it and can replay it. Boris Cherny, who leads Claude Code at Anthropic, summed up the job change, as The New Stack reported (June 2026): he no longer prompts Claude directly, and âmy job is to write loops.â
The loop is also where scale comes from, and where it doesnât. I take that apart in What Running Thousands of Agents Actually Takes. The parts of a harness that coach the model keep getting deleted as models improve; the parts that check its work donât. Thatâs Your Agent Harness Is a List of What the Model Canât Do.
2. Done is a contract, written before the work
The tools that are most explicit about completion all separate a standing bar from per-task criteria. Addy Osmaniâs agent-skills project puts it in one line in its Definition of Done reference: âA task is done only when its acceptance criteria are met and the standing Definition of Done is satisfied.â
The standing part has to stay standing: âA Definition of Done that is renegotiated every sprint is not a Definition of Done.â
In practice thatâs four layers. A floor that no project can lower. A project profile a human ratifies. Acceptance criteria written during planning and frozen when the plan is approved. And per-task verify commands, each mapped back to a criterion.
Two rules from the same projectâs constraint-driven development skill changed how I think about it. âEvery row names the command that produces the verdict.â A target with no command is âan aspiration, not a constraint.â And because agents under pressure âtake the cheapest road to green,â the bar has to be watched: âTightening the bar should be silent; loosening it should be loud.â
This site already has a small version of the floor. A content guard scans every built page for claims Iâve decided never to make, names that shouldnât appear, links to pages that no longer exist, and placeholders that were supposed to be filled. If it finds one, the build fails. Every rule in it is a mistake that already happened once.
The full layering, and how to keep agents from quietly rewriting it, is in Done Is a Contract.
3. A separate judge sees evidence, not claims
The fix for self-praise is structural. In the same Anthropic post: âSeparating the agent doing the work from the agent judging it proves to be a strong lever.â The reason is practical: âtuning a standalone evaluator to be skeptical turns out to be far more tractable than making a generator critical of its own work.â
Separation alone isnât enough, and the post is honest about that too. âOut of the box, Claude is a poor QA agent.â The author watched the evaluator âidentify legitimate issues, then talk itself into deciding they werenât a big deal and approve the work anyway,â and it took several rounds of reading its logs and correcting its prompt before its grading was reasonable.
What makes a judge useful is what itâs given. It gets the artifact, the contract, evidence the harness captured itself (test logs, exit codes, screenshots), and tools to gather more. It doesnât get the builderâs account of why the work is good. Osmaniâs doubt-driven development skill is blunt about why: âIf you hand over conclusions, youâll get back validation of your conclusions.â
Claude Codeâs docs make the same point about context: âA fresh context improves code review since Claude wonât be biased toward code it just wrote.â
The research behind this, and exactly what a validator should and shouldnât see, is in The Agent That Did the Work Shouldnât Decide Itâs Done.
4. Evals measure whether any of it helps
A validator judges one run. It canât tell you whether a prompt change, a model upgrade, or a rewrite of the validator itself made your pipeline better or worse. Only an eval suite can: a fixed set of real tasks, run several times each in clean environments, scored against ground truth the pipeline didnât produce.
The bar to start is lower than people assume. Anthropicâs guide to agent evals (January 2026) says â20-50 simple tasks drawn from real failures is a great start.â Its rule for reading the results is one Iâve adopted: âwe do not take eval scores at face value until someone digs into the details of the eval and reads some transcripts.â
The metric matters as much as the suite. For an agent that runs unattended, the useful number isnât whether it succeeds once in several tries. Itâs whether it succeeds every time. The same guide: âIf your agent has a 75% per-trial success rate and you run 3 trials, the probability of passing all three is (0.75)Âł â 42%.â It recommends pass@k âfor tools where one success matters, pass^k for agents where consistency is essential.â
How to build the suite, and how to measure the validators themselves, is A Validator Is Not an Eval. The arithmetic of unattended reliability is pass^k: The Reliability Math of Unattended Agents.
Earned autonomy, measured
Faberâs README promises âearned autonomy,â and describes the earning as a calendar: conservative on day one, fewer checkpoints by week four, and at month six, â90% autonomous, 10% escalation.â
The research points somewhere better. Autonomy should be earned by measurement, not by elapsed time.
A kind of work moves from assisted to autonomous when its end-to-end pass rate across repeated runs, and its validatorsâ measured catch rate, clear thresholds on its own eval tasks. It moves back when rework in production rises. A calendar measures how long youâve trusted something. It says nothing about whether the trust is warranted.
Thatâs the idea this series is named for, and itâs the same idea behind Freedom Through Autonomy: the business that gives you the most freedom is the one designed to run without you. You only get to walk away from a system youâve measured.
What Iâm not claiming
Most of the numbers in this series come from the labs and tool authors who build these systems and publish what they measure. I cite them because theyâre the best public evidence Iâve found, and I say whose numbers they are.
The evaluator isnât free, and it isnât always worth it. Anthropicâs own conclusion is conditional: a separate evaluator âis worth the cost when the task sits beyond what the current model does reliably solo.â Harnesses arenât permanent either. The same post: âthe space of interesting harness combinations doesnât shrink as models improve. Instead, it moves.â
And Iâm describing a direction for Faber, with the status of each change marked. Some of it has shipped. Most of it is still planned. When I run Faberâs own eval pilot, Iâll publish what it shows, including what didnât work.
What this changes in Faber
The hardening spec is public and its status is âproposed.â Its first principle is the one this series argues for: âCode runs the loop; agents judge inside steps; humans approve irreversible actions.â
Shipped
- A runtime that runs the loop as plain code, with a fresh agent session per step (PR #220).
- Model definitions removed from skills, so skills carry knowledge rather than assumptions about a particular model (PR #233).
Merged, in the next release
- Step verdicts that gate the run: every step returns a structured verdict, the runtime parses it, and a failing or missing verdict stops the run (PR #243).
- Run state saved at every step, so a stopped run can resume where it left off (PR #241).
Planned
- The code-driven runtime for every run, with the chat interface as a thin wrapper that starts it, relays approvals, and summarizes.
- Layered done criteria: project standards kept as data, with organization standards marked as a floor projects can only tighten or a default they can override with a recorded reason.
- Acceptance criteria locked when the plan is approved, so the builder canât rewrite the bar itâs measured against.
- Validators that didnât make the work, running in fresh contexts and reading evidence files the harness writes.
- Task-level state in Build, with verification per task.
- An eval pilot: ten past work items, three runs each, to set the size of the full suite.
- An assumption register recording which model weakness each piece of scaffolding compensates for, re-tested at every model upgrade.
- Autonomy levels tied to measured thresholds instead of a calendar.
Faber is open source on GitHub. The rest of the studioâs stack is on the Platform page, and how the tools fit together is in Building the Fractary Platform.
This is the first post in the Earned Autonomy series. Next: The Agent That Did the Work Shouldnât Decide Itâs Done.