← Blog

Earned Autonomy ¡ Part 1

Who Decides an Agent's Work Is Done? How Faber Is Changing

Agents got good at producing work. Deciding it's done is the hard part now. What the field learned about loops, judges, and evals, and how I'm changing Faber.

Josh McWilliam

9 min read

  • agentic-systems
  • harness
  • evals
  • platform
  • open-source

Getting an agent to produce work is no longer the hard part. Knowing when to believe it is.

Ask an agent whether it finished the job and it will almost always say yes. Anthropic’s engineers put it bluntly in a post on harness design (March 2026): when asked to evaluate work they’ve produced, agents “tend to respond by confidently praising the work—even when, to a human observer, the quality is obviously mediocre.”

So the question I care about most has changed. It used to be “can the agent build this?” Now it’s “who decides it’s done, and on what evidence?”

Why I went looking

Every change in the studio runs through five phases: Frame, Architect, Build, Evaluate, Release. That’s the method the studio’s Faber tooling encodes, and I described it in Vibe Coding Gets You a Demo with one sentence I still stand behind: “The agent that builds a thing is not the one that decides it’s done.”

That’s a principle. I wanted to know what it takes in practice.

So I sat down to revamp Faber and its software pack, faber-code, against what the people building these systems have published: the labs’ engineering write-ups, the documentation for the agent tools I use daily, the open-source spec-driven tools, and the research on why models are bad at grading their own work. The result was a long research note and a hardening spec for Faber.

This series is the readable version. The full research note has every source, and the hardening spec is public on GitHub.

Where Faber started

Faber wasn’t starting from zero. The phases were there. Approvals at the boundaries were there. So was the path from an issue to a pull request, with workflows that inherit from a shared core so each pack only adds what’s specific to its domain.

One important piece had already shipped: a runtime that runs the loop as plain code. fractary-faber workflow-execute walks the plan step by step and starts a fresh agent session for each one (PR #220). No master agent decides what happens next. Code does.

What the research changed is narrower, and it matters more. In most agent tooling, mine included, “done” has meant the steps ran. The goal is for “done” to mean there’s evidence the result works, and that something other than the maker checked it.

Anthropic’s Claude Code documentation describes the default failure in two sentences in its best practices: “Claude stops when the work looks done. Without a check it can run, ‘looks done’ is the only signal available.”

Everything below is about replacing “looks done” with something better.

Four shifts the field converged on

I read a lot of sources that disagree about a lot of things. On these four, they line up.

1. Code owns the loop. The model judges inside each step.

The control-flow question is mostly settled. Deterministic code holds the outer loop: state, stage transitions, budgets, gates, and human sign-off. The model does the judgment inside each step.

Claude Code’s workflow documentation states it plainly: “A workflow script holds the loop, the branching, and the intermediate results itself, so Claude’s context holds only the final answer.” The runtime even makes timestamps and random numbers throw inside a script “so that a relaunched run repeats the same agent() calls.”

The twist is who writes the script. Increasingly it’s the model. The model writes the graph; a deterministic runtime executes it and can replay it. Boris Cherny, who leads Claude Code at Anthropic, summed up the job change, as The New Stack reported (June 2026): he no longer prompts Claude directly, and “my job is to write loops.”

The loop is also where scale comes from, and where it doesn’t. I take that apart in What Running Thousands of Agents Actually Takes. The parts of a harness that coach the model keep getting deleted as models improve; the parts that check its work don’t. That’s Your Agent Harness Is a List of What the Model Can’t Do.

2. Done is a contract, written before the work

The tools that are most explicit about completion all separate a standing bar from per-task criteria. Addy Osmani’s agent-skills project puts it in one line in its Definition of Done reference: “A task is done only when its acceptance criteria are met and the standing Definition of Done is satisfied.”

The standing part has to stay standing: “A Definition of Done that is renegotiated every sprint is not a Definition of Done.”

In practice that’s four layers. A floor that no project can lower. A project profile a human ratifies. Acceptance criteria written during planning and frozen when the plan is approved. And per-task verify commands, each mapped back to a criterion.

Two rules from the same project’s constraint-driven development skill changed how I think about it. “Every row names the command that produces the verdict.” A target with no command is “an aspiration, not a constraint.” And because agents under pressure “take the cheapest road to green,” the bar has to be watched: “Tightening the bar should be silent; loosening it should be loud.”

This site already has a small version of the floor. A content guard scans every built page for claims I’ve decided never to make, names that shouldn’t appear, links to pages that no longer exist, and placeholders that were supposed to be filled. If it finds one, the build fails. Every rule in it is a mistake that already happened once.

The full layering, and how to keep agents from quietly rewriting it, is in Done Is a Contract.

3. A separate judge sees evidence, not claims

The fix for self-praise is structural. In the same Anthropic post: “Separating the agent doing the work from the agent judging it proves to be a strong lever.” The reason is practical: “tuning a standalone evaluator to be skeptical turns out to be far more tractable than making a generator critical of its own work.”

Separation alone isn’t enough, and the post is honest about that too. “Out of the box, Claude is a poor QA agent.” The author watched the evaluator “identify legitimate issues, then talk itself into deciding they weren’t a big deal and approve the work anyway,” and it took several rounds of reading its logs and correcting its prompt before its grading was reasonable.

What makes a judge useful is what it’s given. It gets the artifact, the contract, evidence the harness captured itself (test logs, exit codes, screenshots), and tools to gather more. It doesn’t get the builder’s account of why the work is good. Osmani’s doubt-driven development skill is blunt about why: “If you hand over conclusions, you’ll get back validation of your conclusions.”

Claude Code’s docs make the same point about context: “A fresh context improves code review since Claude won’t be biased toward code it just wrote.”

The research behind this, and exactly what a validator should and shouldn’t see, is in The Agent That Did the Work Shouldn’t Decide It’s Done.

4. Evals measure whether any of it helps

A validator judges one run. It can’t tell you whether a prompt change, a model upgrade, or a rewrite of the validator itself made your pipeline better or worse. Only an eval suite can: a fixed set of real tasks, run several times each in clean environments, scored against ground truth the pipeline didn’t produce.

The bar to start is lower than people assume. Anthropic’s guide to agent evals (January 2026) says “20-50 simple tasks drawn from real failures is a great start.” Its rule for reading the results is one I’ve adopted: “we do not take eval scores at face value until someone digs into the details of the eval and reads some transcripts.”

The metric matters as much as the suite. For an agent that runs unattended, the useful number isn’t whether it succeeds once in several tries. It’s whether it succeeds every time. The same guide: “If your agent has a 75% per-trial success rate and you run 3 trials, the probability of passing all three is (0.75)³ ≈ 42%.” It recommends pass@k “for tools where one success matters, pass^k for agents where consistency is essential.”

How to build the suite, and how to measure the validators themselves, is A Validator Is Not an Eval. The arithmetic of unattended reliability is pass^k: The Reliability Math of Unattended Agents.

Earned autonomy, measured

Faber’s README promises “earned autonomy,” and describes the earning as a calendar: conservative on day one, fewer checkpoints by week four, and at month six, “90% autonomous, 10% escalation.”

The research points somewhere better. Autonomy should be earned by measurement, not by elapsed time.

A kind of work moves from assisted to autonomous when its end-to-end pass rate across repeated runs, and its validators’ measured catch rate, clear thresholds on its own eval tasks. It moves back when rework in production rises. A calendar measures how long you’ve trusted something. It says nothing about whether the trust is warranted.

That’s the idea this series is named for, and it’s the same idea behind Freedom Through Autonomy: the business that gives you the most freedom is the one designed to run without you. You only get to walk away from a system you’ve measured.

What I’m not claiming

Most of the numbers in this series come from the labs and tool authors who build these systems and publish what they measure. I cite them because they’re the best public evidence I’ve found, and I say whose numbers they are.

The evaluator isn’t free, and it isn’t always worth it. Anthropic’s own conclusion is conditional: a separate evaluator “is worth the cost when the task sits beyond what the current model does reliably solo.” Harnesses aren’t permanent either. The same post: “the space of interesting harness combinations doesn’t shrink as models improve. Instead, it moves.”

And I’m describing a direction for Faber, with the status of each change marked. Some of it has shipped. Most of it is still planned. When I run Faber’s own eval pilot, I’ll publish what it shows, including what didn’t work.

What this changes in Faber

The hardening spec is public and its status is “proposed.” Its first principle is the one this series argues for: “Code runs the loop; agents judge inside steps; humans approve irreversible actions.”

Shipped

  • A runtime that runs the loop as plain code, with a fresh agent session per step (PR #220).
  • Model definitions removed from skills, so skills carry knowledge rather than assumptions about a particular model (PR #233).

Merged, in the next release

  • Step verdicts that gate the run: every step returns a structured verdict, the runtime parses it, and a failing or missing verdict stops the run (PR #243).
  • Run state saved at every step, so a stopped run can resume where it left off (PR #241).

Planned

  • The code-driven runtime for every run, with the chat interface as a thin wrapper that starts it, relays approvals, and summarizes.
  • Layered done criteria: project standards kept as data, with organization standards marked as a floor projects can only tighten or a default they can override with a recorded reason.
  • Acceptance criteria locked when the plan is approved, so the builder can’t rewrite the bar it’s measured against.
  • Validators that didn’t make the work, running in fresh contexts and reading evidence files the harness writes.
  • Task-level state in Build, with verification per task.
  • An eval pilot: ten past work items, three runs each, to set the size of the full suite.
  • An assumption register recording which model weakness each piece of scaffolding compensates for, re-tested at every model upgrade.
  • Autonomy levels tied to measured thresholds instead of a calendar.

Faber is open source on GitHub. The rest of the studio’s stack is on the Platform page, and how the tools fit together is in Building the Fractary Platform.


This is the first post in the Earned Autonomy series. Next: The Agent That Did the Work Shouldn’t Decide It’s Done.

The Stack

Build on the stack the studio builds with.

Open-source, Apache-2.0 tools for building agentic systems without vendor lock-in. The same stack behind every venture on this site.