← Blog

Earned Autonomy · Part 3

Done Is a Contract: Four Layers Your Agents Can't Rewrite

An agent stops when the work looks done. A four-layer definition of done, written before the build and enforced by code, is how looks done becomes is done.

Josh McWilliam

9 min read

  • agentic-systems
  • harness
  • production
  • claude-code

An agent’s definition of done is simple: it stopped.

Claude Code’s own best-practices guide says it more politely: “Claude stops when the work looks done. Without a check it can run, ‘looks done’ is the only signal available.” I’ve watched that sentence play out more times than I can count.

The fix isn’t a sterner prompt. It’s a contract: written before the build, kept somewhere the agent can’t quietly edit, and checked by something other than the agent that did the work. The tools that take this most seriously have landed on the same shape, and it has four layers.

Two kinds of done

Addy Osmani’s agent-skills separates them in one line: “A task is done only when its acceptance criteria are met and the standing Definition of Done is satisfied.”

Acceptance criteria are “defined when planning the task” and answer “did we build this thing?” The Definition of Done is “defined once for the project” and answers “is it ready?”

Most agent setups have some version of the first and almost none of the second. Osmani is blunt about why the second has to stay put: “A Definition of Done that is renegotiated every sprint is not a Definition of Done.”

Agents renegotiate all the time. Not out of malice. Out of momentum: when a check is red and the session wants to end, the bar is the cheapest thing in the room to move.

The four layers

The spec-driven tools all separate a standing bar from per-item criteria, each in its own vocabulary.

GitHub’s Spec Kit puts a versioned, ratified constitution above every spec, plan, and task list, and its plan template opens with a constitution check marked “GATE: Must pass before Phase 0 research.” OpenSpec starts from a project config.yaml that injects the stack, the conventions, and per-artifact rules into everything downstream. Anthropic’s long-running harness (November 2025) has an initializer agent write “a structured JSON file with a list of end-to-end feature descriptions” before any coding starts.

Pulled together, the stack looks like this:

LayerWhat it holdsDrafted byApproved byHow often it changes
FloorRules no project can loosen: no new suppressions, skipped or deleted tests, stubs, or secretsThe frameworkIts maintainersRarely, and only tighter
Project profileEvery quality dimension with its command, threshold, direction, and where it runs; exceptions with an owner and an expiryAn agent, from the stack plus a short interviewA named person, by pull requestVersioned
Acceptance criteriaBehaviors with stable IDs, plus an explicit out-of-scope listThe planning stepA human, at the planning gateFrozen once approved
Task checksA few acceptance bullets per task and the command that proves them, mapped back to criteria IDsThe design stepA coverage check in codeEvery task

Read it top to bottom and the higher a layer sits, the slower it changes and the more deliberate its approval. The floor almost never moves. The task checks change every task.

None of it is approved by the agent doing the building.

Planning writes the criteria, at the right altitude

Planning matters because it’s where criteria come from. In Anthropic’s March 2026 write-up on harness design, the planner survived every simplification for a plain reason: “Without the planner, the generator under-scoped.”

But the planner was kept at product level, deliberately. The worry was that “if the planner tried to specify granular technical details upfront and got something wrong, the errors in the spec would cascade into the downstream implementation.”

So I’d split it this way. The planning step owns what must be true: acceptance criteria with stable IDs and a list of what’s out of scope. The design step owns how each item will be proven: a check per task, and a coverage map back to the criteria.

That coverage map is mechanical, which means code can check it. Spec Kit’s analyze command is “STRICTLY READ-ONLY” and flags a “requirement with zero coverage that blocks baseline functionality” as critical.

There’s an optional fifth layer for work at the edge of what a model can do. In the same Anthropic harness, “the generator and evaluator negotiated a sprint contract: agreeing on what ‘done’ looked like for that chunk of work before any code was written.” The generator “proposed what it would build and how success would be verified,” and the evaluator reviewed the proposal. The contracts were granular: “Sprint 3 alone had 27 criteria covering the level editor.”

When a newer model arrived, the author removed the sprint construct entirely. That tells you which part is scaffolding. The contract step depends on the model. The layers don’t.

The weak point: the worker ticks its own boxes

Here’s the gap I’d look for in any spec-driven setup, including my own.

Spec Kit’s implement command tells the implementing agent: “For completed tasks, make sure to mark the task off as [X] in the tasks file.” Its checklist gate is real, and when items are unchecked it asks, “Do you want to proceed with implementation anyway? (yes/no)”. OpenSpec’s command reference says of its verify step: “Does not block archive, but surfaces issues.”

That’s a reasonable default for interactive work, where a person is reading along. Unattended, a box the worker ticks is just a box the worker ticked.

Only two things take the decision away from the worker: a deterministic check, or a verdict from a separate judge. The judge is the previous post. This one is about the contract the judge reads.

Make the bar hard to lower quietly

The sharpest line on this comes from the constraint-driven development skill in agent-skills: “Agents don’t craft clever loopholes. They hit a red check and take the cheapest road to green.”

It lists five moves to watch for in the diff: the threshold moved, a test got easier, a checker got silenced, work is unfinished, an exception appeared. None of them needs more than git diff to catch. And then the rule I’ve adopted wholesale: “Tightening the bar should be silent; loosening it should be loud.”

Three mechanisms make that real:

  1. Keep status in JSON. Anthropic’s November harness moved its feature list to JSON because “the model is less likely to inappropriately change or overwrite JSON files compared to Markdown files.”
  2. Make gap-filling append-only. Spec Kit’s converge command is “APPEND-ONLY, NEVER REWRITE”: it can add a task for unbuilt work, but it can’t rewrite, renumber, or delete one.
  3. Diff the bar against the branch point. Compare the constraints file to its state where the branch started. A weaker threshold is a finding, not a change.

The corollary: once the criteria are approved, lock them. If the build needs different criteria, that’s a planning question, and it goes back to planning with a person in the loop.

Every row names a command

A bar you can’t run is a wish. The same skill says it better: “Every row names the command that produces the verdict. A dimension with a number and no command in this column is an aspiration, not a constraint.”

Three refinements make a profile honest:

Ratchets instead of invented targets. “Set 80% coverage on a codebase at 62% and you get a red build forever, then a team that learns to ignore red builds.” Record today’s number and refuse to get worse.

Applicability, said out loud. “Lighthouse and axe need a URL.” A command-line tool or a library has no URL to hit, so the profile drops those dimensions explicitly rather than faking them.

At least one check the agent can’t argue with. The skill ranks checks by one question: can the agent make this pass by writing code that doesn’t work? External tools such as axe-core or osv-scanner sit at the top (“The agent can’t argue with these”). The project’s own test suite sits at the bottom, “the only genuinely circular one.” Its instruction: “Check that at least one external constraint is present.”

One entrypoint, every boundary

The last piece is where the contract gets enforced, and the answer is everywhere, with the same command.

A community extension in Spec Kit’s catalog describes the goal in one line: “One policy file, one verify entrypoint, identical results at every boundary.” The boundaries are the agent’s hooks, the git hooks, and CI.

Same command doesn’t mean same depth. Agent-skills splits checks into a fast tier after each edit, a task tier when the agent thinks it’s done, and a full tier in CI. The reason: “A check that stalls the agent gets switched off, and a gate people switched off is worse than no gate, because the bar still looks like it exists.”

Then there’s the exit code. The skill’s floor guard returns 0 for clean, 1 for a violation, and 2 when the guard could not run at all. Its warning is the most important sentence in this post: “Never let a 2 read as a 0.” A check that didn’t run is not a check that passed.

Claude Code gives you the hooks to wire this in. A Stop, SubagentStop, or TaskCompleted hook that exits with code 2 blocks the action; for TaskCompleted, the hooks reference says it “Prevents the task from being marked as completed.”

There’s a limit, and it’s a good one: “after stop hooks have continued the turn eight times in a row, Claude Code overrides the next block and ends the turn.” A gate the agent can’t satisfy shouldn’t loop. It should escalate to a person.

For an organization, managed settings with allowManagedHooksOnly restrict which hooks run, so a floor shipped centrally can’t be quietly switched off one project at a time.

What this looks like on this site

This site has a floor, and I’ve written about it before. A content guard reads every built page and fails the build on about two dozen rules: names I don’t use, claims I don’t make, links to pages that no longer exist, placeholders that were supposed to be filled in.

Every rule is a mistake that already happened once. That’s how a floor should grow. It isn’t an external check in the skill’s sense: it lives in the repo, and an agent could edit it. But it’s a plain string match, so arguing with it gets an agent nowhere, and any change to the guard itself shows up in the diff I review.

It has one entrypoint. A single command runs the type check, the full build, and the guard, and fails if any of them fails. The agent runs the same command I would, until it passes. Then I read the pull request.

The acceptance criteria live in the spec. When this site was repositioned in September, the spec had a decision table: what’s live, what’s coming soon, and what’s out of scope. The build log from that week walks through it.

What this changes in Faber

Faber runs work for many projects across several packs: software, content, data ingests, cloud infrastructure. A validator that’s rewritten for every project isn’t a validator; it’s a fork. The project-specific part has to be data, read by one generic engine.

Here’s what I’m changing, all of it planned in Faber’s hardening spec:

  • Project standards as data. Each project gets a standards profile that its maintainers own and change by pull request. Organization standards arrive through Codex, pinned by version and marked either floor (a project may only tighten it) or default (a project may override it, with a recorded reason).
  • Criteria locked at approval. Acceptance criteria and task definitions are written during planning and locked when they’re approved. The agent doing the build writes only status and evidence; changing a criterion sends the work item back through planning.
  • One way to check a criterion, many kinds of check. Each criterion says how it’s verified: a command, a metric, a schema, a rubric, a human, or a trial run. One validator reads the profile and runs whichever kind each criterion names.
  • One verify entrypoint, fractary-faber verify --stage, called by the runtime, by hooks, and by CI, writing the evidence files that completion requires. Exit 0 passes, 1 fails, and 2 means “could not run,” which never counts as a pass.

The foundation is already there: Faber’s command-line runtime runs the loop as plain code, with a fresh agent session for each step. The contract is what that loop will check against.


This is the third post in the Earned Autonomy series. Previous: The Agent That Did the Work Shouldn’t Decide It’s Done. Next: What Running Thousands of Agents Actually Takes.

The Stack

Build on the stack the studio builds with.

Open-source, Apache-2.0 tools for building agentic systems without vendor lock-in. The same stack behind every venture on this site.