← Blog

Earned Autonomy · Part 8

Agent Harnesses and Evals: The Full Research Note

The full research note behind the Earned Autonomy series: how code owns the loop, a separate judge decides done, and evals show whether any of it helps.

Josh McWilliam

35 min read

  • agentic-systems
  • harness
  • evals
  • claude-code
  • production

This is the research note behind the Earned Autonomy series. I wrote it while revamping Faber, the open-source workflow tool the studio builds with, to match what the field has learned about running agents unattended. Every claim links to its source.

It’s long and dense on purpose. If you want the arguments one at a time, the series posts are the readable version:

  1. Who Decides an Agent’s Work Is Done? How Faber Is Changing
  2. The Agent That Did the Work Shouldn’t Decide It’s Done
  3. Done Is a Contract: Four Layers Your Agents Can’t Rewrite
  4. What Running Thousands of Agents Actually Takes
  5. Your Agent Harness Is a List of What the Model Can’t Do
  6. A Validator Is Not an Eval: How to Measure the Judge
  7. pass^k: The Reliability Math of Unattended Agents

The note reflects sources published up to early October 2026, and parts of it will date. The Faber changes it recommends are tracked in the public harness hardening spec.

Script the loop and separate the judge

Leading practice for autonomous software development has converged on one architecture. Deterministic code owns the outer loop: state, stage transitions, budgets, gates and human sign-off. The model owns judgment inside each step, and increasingly writes task-specific orchestration scripts that a runtime replays deterministically. And “done” is decided by evidence and a separate judge rather than by the agent that did the work.

For Faber, four shifts follow.

  1. The definition of done should be a layered, file-based contract that one generic validator interprets, instead of validators rewritten per project: a framework floor, a human-ratified project profile, acceptance criteria drafted in Frame and frozen at the Frame→Architect gate, and per-task verify commands written in Architect.
  2. Validators should run as fresh-context agents that receive the artifact, the contract and harness-captured evidence, but never the implementer’s reasoning or claims. Self-review is systematically lenient and shared context feeds reward hacking. Separation pays off most when the judge can execute checks itself.
  3. “Don’t build agents, build skills” rejected bespoke per-domain agent architectures, not agent instances. Anthropic shipped workflows that run hundreds of subagents within months of that talk, and Boris Cherny’s “thousands” of nightly agents are machine-spawned runs fenced by cheap verification.
  4. Per-step validators are runtime gates, not evals. Only an offline suite of real tasks, run several times each in clean environments and scored on pass^k, can say whether a validator, prompt or model change helped, and whether the validators catch what they claim to catch.

Done is a four-layer contract that planning writes and code enforces

The tools that spell out completion most explicitly all separate a standing bar from per-item criteria. Addy Osmani’s agent-skills puts it in one line: “A task is done only when its acceptance criteria are met and the standing Definition of Done is satisfied,” where acceptance criteria are “defined when planning the task” and the Definition of Done is “defined once for the project” and reused, because “a Definition of Done that is renegotiated every sprint is not a Definition of Done” (agent-skills, definition-of-done.md).

GitHub Spec Kit encodes the same stack as constitution → spec → plan → tasks, with a semver-versioned, ratified constitution (Spec Kit constitution command) and a plan-stage “Constitution Check” marked “GATE: Must pass before Phase 0 research” (Spec Kit plan template). OpenSpec uses project config.yaml → requirement specs → change deltas → tasks (OpenSpec customization). Anthropic’s long-running harnesses compress the lower layers into a JSON feature list written by an initializer agent (Anthropic, November 2025) and add a per-chunk “sprint contract” that the generator and evaluator negotiate “before any code was written” (Anthropic, March 2026). The evidence for the layering is the tools’ own documentation, and the pattern holds across all of them.

LayerHoldsDrafted byRatified byFaber home
L0 floorStack-agnostic rules: no new suppressions, skipped or deleted tests, stubs or secrets; the bar may not be weakenedFrameworkFramework maintainers; shipped as a managed hookIdentical in every project
L1 project profilePer dimension: command, threshold, direction, stage, applicability; exceptions with an owner and an expiryAgent, from stack detection plus a short interviewA named human; versionedProject config read by every phase
L2 acceptance criteriaStable-ID behaviors (EARS, Given/When/Then, or SHALL plus scenarios) and an out-of-scope listFrame agentHuman, or autonomy policy, at the Frame→Architect gateFrame output, hash-locked
L3 task verificationPer-task acceptance (three bullets at most) plus a verify command, mapped to L2 IDsArchitect agentDeterministic coverage checkArchitect output
L2.5 contract (optional)Granular testable behaviors for one chunk of BuildImplementer proposesValidator approves; may only tighten L0–L2Start of Build

Planning phases matter because they are where criteria come from, but they should stop at the right altitude. Anthropic kept its planner through every simplification because “without the planner, the generator under-scoped,” yet confined it to product-level specs because “if the planner tried to specify granular technical details upfront and got something wrong, the errors in the spec would cascade into the downstream implementation” (Anthropic, March 2026).

The sprint contract bridged that gap: the generator “proposed what it would build and how success would be verified, and the evaluator reviewed that proposal,” and one sprint alone carried 27 criteria, granular enough that the evaluator’s failures pointed at specific handlers and routes.

For Faber, Frame should own what must be true (L2, testable and ID-stamped). Architect should own how each criterion will be proven (L3, with a coverage map back to L2 in the style of Spec Kit’s strictly read-only /analyze, which flags uncovered requirements that block baseline functionality as critical (Spec Kit analyze)). Build can open with a contract step for work at the edge of model capability. Anthropic dropped the sprint construct when a newer model arrived, so the contract step is model-dependent scaffolding; the layers themselves are not.

The weak point: the implementer still marks its own completion

Spec Kit’s implementer ticks [X] in its own task file, and its checklist gate can be waved through with a “yes” (Spec Kit implement). OpenSpec’s verify step “does not block archive, but surfaces issues” (OpenSpec commands). Only an independent evaluator’s verdict or a deterministic guard takes the decision away from the worker.

The mechanisms that stop criteria being quietly weakened are concrete:

  • Keep status in JSON, because “the model is less likely to inappropriately change or overwrite JSON files compared to Markdown files” (Anthropic, November 2025).
  • Make post-build gap-filling append-only, as Spec Kit’s /converge does (“APPEND-ONLY, NEVER REWRITE”) (Spec Kit converge).
  • Diff the bar against the branch point, as agent-skills’ floor guard does, because “agents don’t craft clever loopholes. They hit a red check and take the cheapest road to green,” so “tightening the bar should be silent; loosening it should be loud” (agent-skills CDD skill).

The same skill ranks checks by circularity, from external tools such as axe or osv-scanner (“the agent can’t argue with these”) down to the project’s own test suite (“the only genuinely circular one”), and requires at least one external check per project.

Faber should hash-lock L1 and L2 at their approval gates, fail Evaluate closed if either changed without re-approval, give Build write access only to a status-and-evidence file, and route any change to the criteria back to Frame.

Validators stay generic when the project-specific part is data

The converging recipe is a single engine plus a per-project declaration in which every dimension names its command, threshold, direction (minimum, maximum or ratchet), stage and applicability. agent-skills insists that “every row names the command that produces the verdict,” because “a dimension with a number and no command 
 is an aspiration, not a constraint” (agent-skills CDD skill).

The engine detects the stack before asking questions, replaces invented targets with ratchets (“Set 80% coverage on a codebase at 62% and you get a red build forever”), marks inapplicable dimensions explicitly (Lighthouse and axe “need a URL,” so a CLI or library drops them rather than faking them), and scopes exceptions by path with an owner and an expiry. Customization flows through layered overrides, such as Spec Kit’s priority-ordered preset stack (Spec Kit presets) and Claude Code’s hook scopes, not through forked validators.

In Faber terms, the validator becomes an interpreter that dispatches on each criterion’s verification method (test, command, end-to-end, HTTP, snapshot, policy, evaluator or manual) and reads commands from the project profile. Web, library and CLI presets follow directly from the sources; infrastructure-as-code and data-pipeline presets are reasonable extrapolations that none of the surveyed tools ships yet.

Enforcement has to live below the prompt

The robust pattern is one verify entrypoint run at every boundary. Spec Kit’s community gates extension promises “one policy file, one verify entrypoint, identical results at every boundary” across agent hooks, git hooks and CI (Spec Kit community catalog), and agent-skills splits the same checks into fast, task and full tiers, warning that “a check that stalls the agent gets switched off.”

Claude Code’s Stop, SubagentStop and TaskCompleted hooks can refuse completion with exit code 2, but the harness overrides a stop hook after eight consecutive forced continuations, so a gate that cannot be satisfied must escalate rather than loop. Managed-policy settings with allowManagedHooksOnly let an organization ship a floor that projects cannot silently disable (Claude Code hooks).

A fractary-faber verify --stage command with a 0/1/2 exit contract, where 2 means “could not run” and never counts as a pass (“Never let a 2 read as a 0”) (agent-skills floor guard), wired into hooks, the executor and CI, would give Faber the same property.

Self-review fails exactly where validators matter most

Anthropic’s own harness work states the problem plainly: “When asked to evaluate work they’ve produced, agents tend to respond by confidently praising the work—even when, to a human observer, the quality is obviously mediocre.” The fix was structural, not rhetorical: “Separating the agent doing the work from the agent judging it proves to be a strong lever,” because “tuning a standalone evaluator to be skeptical turns out to be far more tractable than making a generator critical of its own work, and once that external feedback exists, the generator has something concrete to iterate against” (Anthropic, March 2026).

Claude Code’s documentation gives the incentive version of the same argument: “Claude stops when the work looks done. Without a check it can run, ‘looks done’ is the only signal available,” and “a fresh context improves code review since Claude won’t be biased toward code it just wrote” (Claude Code best practices).

The research literature explains why, through four reinforcing mechanisms. (The paper figures below are from abstracts unless marked as from the full text.)

Models are worse at finding their own errors than at fixing located ones. Intrinsic self-correction without external feedback struggles and “at times” degrades results (Huang et al., ICLR 2024). A TACL survey found self-correction works well in tasks that can use reliable external feedback, while no prior work showed it succeeding with feedback from prompted LLMs outside tasks exceptionally suited to it (Kamoi et al., 2024). Models backtrack successfully once told where the mistake is (Tyen et al., 2024). Code self-repair is “bottlenecked by the model’s ability to provide feedback on its own code,” while a stronger model’s feedback helps far more (Olausson et al., ICLR 2024). And models fix errors in user input but miss identical errors in their own output, a 64.5% average “blind spot” across 14 open models (Self-Correction Bench, 2025).

Judges prefer what they recognize. One frontier model was 73.5% accurate at telling its own outputs from those of two other LLMs and humans (full text), and self-preference rises linearly with self-recognition (Panickssery et al., NeurIPS 2024), while low-perplexity, familiar text gets higher scores than human evaluators give it, whoever wrote it (Wataoka et al., 2024). Crucially, harmful self-preference persists in the cases where the evaluator erred as the generator, and stronger models show more of it when they err (Chen et al., 2025), so self-review is weakest precisely on the defects a validator exists to catch.

Shared context breeds reward hacking. When one model generated and evaluated in the same context window, evaluator ratings rose while human-judged quality stagnated or fell, and severity depended on how much context the two roles shared (Pan et al., 2024).

The agent that wants to stop is the one deciding it is done. Anthropic watched its evaluator “identify legitimate issues, then talk itself into deciding they weren’t a big deal and approve the work anyway.”

Separation is necessary but not sufficient

“Out of the box, Claude is a poor QA agent,” and Anthropic’s separate evaluator also tested superficially until the author spent “several rounds” reading its logs and correcting its prompt against his own judgment (Anthropic, March 2026).

The large measured gains come from judges with an external signal. An agentic judge that inspects the workspace reached roughly 90% alignment with the human judges’ consensus versus about 70% for a plain LLM judge on a 55-task development benchmark (full text) (Agent-as-a-Judge, 2024), and Anthropic’s evaluator drove the running app through Playwright and returned file:line causes rather than opinions.

Anthropic’s post names a single model for the whole harness, so its separation was of context, prompt and tools rather than model. A fresh context removes the shared-context factor; only a different model family or a tool-grounded check addresses familiarity bias, which is why panels of judges from disjoint model families beat a single large judge at over seven times lower cost (Verga et al., 2024). The case for separation rests on evidence converging from several directions, which is strong enough to act on.

Where Faber started

Faber’s phases, approvals and issue-to-pull-request flow were already in place. The hardening spec records where the work on “done” begins: in plugin mode, every step, validators included, ran in one shared context, and in the CLI runtime, which already gives each step a fresh agent session, a step’s verdict didn’t yet gate the run (issue #235). That fix, verdicts that stop a run when they fail, has since merged and goes out in the next release (PR #243), along with saved run state and resume (PR #241). Moving validators into fresh contexts only helps once their verdicts gate, so parsing a structured verdict and failing closed comes first.

What a validator should see, and what it must never see

The sources agree on a four-part input packet and one exclusion. The validator gets:

  • the artifact: diff, changed paths, commit, and how to build or run it;
  • the contract: L1–L3 criteria, the out-of-scope list, project rules and severity definitions;
  • evidence the harness captured itself: raw test logs, exit codes, screenshots;
  • tools to gather its own: read-only repository access, a test runner, a browser.

It does not get the claim. Osmani’s doubt-driven-development skill is blunt: “A fresh-context reviewer needs the artifact and the contract, not the journey
 Strip your reasoning. If you hand over conclusions, you’ll get back validation of your conclusions” (agent-skills doubt-driven-development). Claude Code’s docs describe the same boundary: a fresh-context reviewer “sees only the diff and the criteria you give it, not the reasoning that produced the change, so it evaluates the result on its own terms” (Claude Code best practices).

Anthropic’s official code-review plugin does pass each reviewer the pull request’s title and description, to “provide context regarding the author’s intent” (claude-code code-review command). The two positions reconcile if Faber passes intent but scrubs assertions such as “all tests pass” or “I verified X” from the handoff. The rule of thumb: give “what should be true” and withhold “why the author believes it is true.”

Evidence must be produced, not reported. Claude Code’s /goal evaluator illustrates the limit of transcript-only judging: it “doesn’t run commands or read files independently,” so it can only be as good as what the worker chose to surface (Claude Code /goal). A gating validator should either run the checks itself or read evidence files the harness wrote, following the example verification rule in Anthropic’s managed Code Review docs, “behavior claims need a file:line citation in the source, not an inference from naming” (Claude Code Code Review).

Missing context is a common source of false positives, and the sources offer two designs: a diff-only reviewer told “do not flag issues that you cannot validate without looking at context outside of the git diff” (claude-code code-review command), which is cheaper with fewer findings, or full-codebase access plus a per-finding verification pass, which costs more and catches more (Claude Code Code Review). Faber can pick by risk tier.

Output should be per-criterion verdicts with severity, an explicit “Unknown” escape hatch that Anthropic recommends for judges (Anthropic, January 2026), and an “unverified” state that is never counted as a pass, mirroring how Claude Code workflows report claims their verifiers could not check (Claude Code workflows). Scope matters too: “A reviewer prompted to find gaps will usually report some, even when the work is sound,” and chasing every finding produces over-engineering, so validators should flag only gaps against correctness and the stated requirements (Claude Code best practices).

Two to three review cycles, then escalate or decompose

Practitioner bounds cluster at two to three LLM review cycles per artifact. Doubt-driven development stops after three (“escalate to user, don’t grind a fourth alone”) and treats a need for more as a sizing failure: “the artifact is too big — return to Step 2 and decompose. Do not lift the bound.” It also names “doubt theater,” where repeated cycles surface findings that are never acted on, as a signal to stop (agent-skills doubt-driven-development).

Claude Code advises clearing context “after two failed corrections” and re-prompting with what was learned (Claude Code best practices), and its managed reviewer suggests suppressing new nits after the first review so a one-line fix does not reach “round seven on style alone.” Anthropic’s subjective-design loop ran 5 to 15 iterations but plateaued, and the author “regularly saw cases where I preferred a middle iteration over the last one,” which argues for keeping the best-scoring checkpoint rather than the last (Anthropic, March 2026).

Faber’s spec caps fix cycles at three, then escalates or splits the work, which sits inside this band. Adding a no-progress stop (same finding recurring, unchanged artifact) and a fresh implementer context after two failed fixes would close the remaining loopholes.

The evaluator is cheap relative to what it protects

In Anthropic’s simplified harness, the three QA rounds cost $10.39 of a $124.70 run (about 8%) and 25 of roughly 230 minutes, while the build rounds dominated spend. QA still caught that “audio recording is still stub-only” and that core features were “display-only” (Anthropic, March 2026). Anthropic’s conclusion is conditional: the evaluator “is worth the cost when the task sits beyond what the current model does reliably solo.”

Its managed PR reviewer costs $15–25 per review and raised the share of PRs receiving substantive comments from 16% to 54%, with under 1% of findings marked incorrect by engineers, a vendor-reported precision figure without a recall measurement (Claude blog, Code Review).

Fresh contexts pay off for reads, verification and parallel transforms, not coupled writes

Validators are not the only steps that benefit from isolation. Work belongs in a separate context when it is read-heavy (exploration, search, audit); when it produces verbose output that only needs a summary (subagents typically return a “condensed, distilled summary
 often 1,000-2,000 tokens” (Anthropic, September 2025)); when it is independent verification; when it is a per-item transform with an objective per-item check (file-disjoint, in worktrees); or when it needs narrower permissions, such as a read-only reviewer.

Work belongs in one context, or with a single writer, when decisions are tightly coupled, when reasoning is sequential, or when phases need constant back-and-forth. Cognition, whose June 2025 “Don’t Build Multi-Agents” was the field’s main counterweight, revisited it in April 2026: “multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions. A clean-context reviewer catches bugs the coder can’t see” (Cognition, April 2026). A study of 260 agent configurations across six benchmarks found multi-agent performance ranging “from +80.8% on decomposable financial reasoning to -70.0% on sequential planning,” and that “architectures without centralized verification tend to propagate errors more than those with centralized coordination” (Kim et al., December 2025). Anthropic’s own multi-agent research system beat a single agent by 90.2% but used about 15× the tokens of chat, and was explicitly a poor fit for “most coding tasks,” which have fewer truly parallelizable parts (Anthropic, June 2025).

Mapped onto Faber, this gives each phase a shape:

  • Frame fans out read-only exploration and has one synthesizer write the acceptance criteria.
  • Architect can generate competing designs in parallel, but one decider records the decision in an artifact every downstream step receives in full, honoring Cognition’s rule to share context rather than fragment it.
  • Build keeps one writer per module, parallelizing only across partitioned, independently testable units in separate worktrees, with an integration gate on the full suite.
  • Evaluate runs fresh-context validators and, for high-stakes findings, adversarial refuters.

Across all phases, a fresh session per step suits long unattended runs, because subagents and fresh sessions inherit no conversation history (Claude Code subagents). The phase artifacts, not the chat transcript, must carry everything the next step needs.

Skills, agents and workflows are layers, not rivals

“Don’t build agents, build skills” was an argument about architecture, not about how many agents to run. Barry Zhang and Mahesh Murag made it at the AI Engineer Code conference in November 2025 (AI Engineer talk page), and Anthropic’s January 22, 2026 write-up spells out what was abandoned: “We used to think agents in different domains would look very different. A coding agent, a research agent, one for finance, one for marketing—each seemed to need its own tools and scaffolding.” Instead, “Claude Code is a coding agent, but also a general-purpose agent that happens to work through code,” and skills “turn a capable generalist into a knowledgeable specialist.”

The post’s summary of the stack is the clearest statement of the layering: “the loop reasons, the runtime executes, MCP connects, and skills guide” (Claude blog, January 2026). What the industry moved away from is the bespoke agent class per domain, each with hand-built scaffolding. What it kept, and multiplied, is agent instances running one general harness.

Anthropic’s documentation treats the pieces as composable layers. Skills hold procedural knowledge, subagents handle task delegation and MCP handles tool connectivity, and “a code-review subagent can use Skills for language-specific best practices” (Claude blog, Skills explained). In Claude Code a subagent can preload skills, a skill marked context: fork runs inside a fresh subagent (Claude Code skills), and the bundled /batch command is itself a skill that splits a change across 5 to 30 worktree-isolated subagents (Claude Code, run agents in parallel). Managed Agents defines an agent as “the model, system prompt, tools, MCP servers, and skills” (Managed Agents overview).

Barry Zhang, credited on the January skills post, also co-wrote the multi-agent research post, and the company launched dynamic workflows on May 28, 2026, in which “Claude dynamically writes orchestration scripts that run tens to hundreds of parallel subagents in a single session, checking its work before anything reaches you” (Claude blog, May 2026). The docs now sort the primitives by who holds the plan: Claude turn by turn (subagents and skills), a lead agent (agent teams), or a script (workflows) (Claude Code workflows). A companion page chooses among them by three questions: who coordinates, whether workers need to talk, and whether they touch the same files, and still marks agent teams experimental (Claude Code, run agents in parallel).

Cherny’s thousands are machine-spawned runs fenced by verification

Boris Cherny’s public numbers climbed over a few months, and most of them reach us through press coverage of talks rather than transcripts.

  • In January, coverage of his X thread described 5 terminal sessions plus 5–10 web sessions, and his advice to give Claude a way to verify its work, which he argued improves final quality “2-3x” (VentureBeat, January 2026). That multiplier is an informal estimate. Anthropic’s current docs keep the advice and drop the number: a check Claude can run is “the difference between a session you watch and one you walk away from” (Claude Code best practices).
  • In a Sequoia Capital interview in May he still ran “five to 10 sessions,” but “usually, every night, I have like a few thousand that are doing kind of deeper work” (Business Insider, May 2026).
  • In June, The New Stack reported that he no longer prompts Claude directly and that “my job is to write loops” (The New Stack, June 2026).
  • At Fortune Brainstorm Tech on June 8: “This morning I was managing maybe a few hundred
 Some days it’s 
 thousands, or tens of thousands,” and he had not written code by hand in about eight months (Fortune, June 2026).
  • At Meta’s @Scale he described a transition “to the point where agents are prompting agents that then write the code,” called loops “just as important and as big a step” as the move from hand-written code to agents, and gave examples of standing loops that hunt for architecture improvements and duplicated abstractions (TechCrunch, June 2026).

A widely shared line, “you build the harness that runs the loops,” has no primary source; it circulated through a third-party post and reads as a paraphrase of “my job is to write loops.”

Read together, the numbers say less about human attention than about automation. The sessions Cherny supervises stayed at roughly 5–10; the hundreds and thousands are subagent runs inside workflows, loops and routines. Every large-scale example on record rests on a cheap, objective oracle.

  • Nicholas Carlini’s 16 parallel agents ran about 2,000 sessions over two weeks for just under $20,000 to produce a 100,000-line C compiler that builds Linux 6.9, with GCC as a differential oracle and the warning that “it’s important that the task verifier is nearly perfect” (Anthropic, February 2026).
  • Bun’s Zig-to-Rust rewrite ran 64 agents (four workflow shards of 16), with the compiler’s error output as the work queue and the test suite as the bar (The Pragmatic Engineer, July 2026).
  • Workflows default to 16 concurrent agents, configurable up to 256, with a hard cap of 1,000 per run (Claude Code workflows).

Scale is the product of scripted orchestration and machine verification, not of more supervision.

For Faber, the lesson is to put knowledge in skills and keep agents generic. Phase procedures, rubrics, the done-criteria interpreter’s instructions and each validator’s criteria belong in versioned, progressively disclosed skills that can be preloaded into a validator subagent. Agent definitions should stay thin and exist mainly to scope permissions, such as a read-only reviewer. A validator then becomes “generic subagent + validator skill + fresh context + schema-validated output,” not a hand-built agent class per phase.

Anthropic’s skill-creator guidance adds a durability test: “capability uplift” skills become obsolete once the base model “starts passing your evals without the skill loaded,” while “encoded preference” skills stay useful (Claude blog, skill-creator). Faber’s project-specific done profiles are the durable kind; its generic “how to code well” prompts are the perishable kind and should be re-tested at every model upgrade.

Harnesses are thinning while verification and infrastructure thicken

The control-flow question is largely settled, with one twist. Claude Code’s workflow documentation states the consensus: “A workflow moves the plan into code. With subagents, skills, and agent teams, Claude is the orchestrator: it decides turn by turn what to spawn or assign next, and every result goes into a context window. A workflow script holds the loop, the branching, and the intermediate results itself, so Claude’s context holds only the final answer.” The runtime even makes Date.now() and Math.random() throw inside scripts “so that a relaunched run repeats the same agent() calls” (Claude Code workflows).

OpenAI’s Agents SDK docs describe code-based orchestration as more deterministic and predictable than LLM-based routing (OpenAI Agents SDK). HumanLayer’s 12-Factor Agents makes “own your control flow” a principle so that loops can pause for humans and resume durably (12-Factor Agents). And the long-run case studies, from Geoffrey Huntley’s Ralph loop to Anthropic’s initializer harness and Carlini’s compiler fleet, keep the outer loop in bash or code and state in files and git (how-to-ralph-wiggum).

The counter-evidence is narrower than it looks. LLM lead agents work for breadth-first research; Cursor found that planners, workers and a judge agent outperformed flat self-coordination at fleet scale (Cursor, January 2026); and Anthropic’s April 2026 guidance tells builders to “let Claude orchestrate its own actions” through code execution rather than hand-written tool plumbing (Claude blog, April 2026).

The twist is that the model now writes the orchestration: the model writes the graph, and a deterministic runtime executes and replays it. “Graph engineering” is not an established term; the field says workflows versus agents, code versus LLM orchestration, dynamic workflows, harness engineering (popularized by OpenAI’s February 2026 post and Birgitta Böckeler’s articles) and, most recently, loop engineering.

Components that compensate for model weaknesses keep going stale

Anthropic’s principle is that “every component in a harness encodes an assumption about what the model can’t do on its own, and those assumptions are worth stress testing
 because they can quickly go stale as models improve.” Its own harness dropped context resets once a newer model stopped wrapping up early near its context limit, then dropped sprint decomposition with the generation after that and moved evaluation to a single end-of-run pass. A radical cut failed, so the method that worked was removing one component at a time (Anthropic, March 2026).

The conclusion was not that harnesses vanish: “the space of interesting harness combinations doesn’t shrink as models improve. Instead, it moves.” Minimal agents make the same point from the other side: mini-swe-agent is about 100 lines of Python using only bash and scores above 74% on SWE-bench Verified (mini-swe-agent).

What does not go stale is governance and infrastructure: oracles and verification, security boundaries, durable state and resume, budget caps, isolation and approval gates. Böckeler’s split of a harness into feedforward “guides” and feedback “sensors” adds a sobering limit: neither deterministic nor LLM-based sensors reliably catch “misdiagnosis of issues, overengineering and unnecessary features, misunderstood instructions” (Böckeler, martinfowler.com), which is the strongest argument for keeping a human at the framing gate and at outcome review.

Five directions

  1. Harnesses are becoming per-task and model-authored. With ultracode, “a single request can turn into several workflows in a row: one to understand the code, one to make the change, and one to verify it” (Claude Code workflows).
  2. Execution is moving onto managed runtimes that separate the “brain,” the sandboxed “hands” and a durable session log outside the context window, with credentials “never reachable from the sandbox where Claude’s generated code runs” (Anthropic, Managed Agents, April 2026). Managed Agents also offers “outcomes,” which define done through a rubric and a grader that runs in a separate context window, a hosted version of the validator pattern (Managed Agents, define outcomes).
  3. Agents are becoming event-driven. Routines fire on schedules, API calls or GitHub events; their fired prompts “can’t act as approval or consent,” and “a green status
 does not mean the task in your prompt succeeded” (Claude Code routines).
  4. Effort is shifting to the environment. OpenAI’s harness-engineering account describes about a million lines and roughly 1,500 PRs with zero hand-written code, a short AGENTS.md acting as a table of contents, and custom linters whose error messages carry remediation instructions for agents (OpenAI, February 2026).
  5. Verification is the bottleneck. Carlini warns that “the thought of programmers deploying software they’ve never personally verified is a real concern” (Anthropic, February 2026), and the Agent SDK guidance ranks rule-based feedback first while calling LLM-as-judge “generally not a very robust method” (Claude blog, September 2025).

The named patterns, as step types

PatternUse it whenFaber step type
Loop until a check passes (evaluator-optimizer)Objective criteria exist: type check, tests, rubricVerify→fix loop with a cap and a “two rounds without progress” stop
Fresh-session task loop (Ralph, initializer/coder)Long builds with backpressure from tests, types and buildBuild as one task per session, JSON task state, one commit per task
Fan-out and synthesizeMany independent items or sourcesParallel read-only step with one synthesizer (Frame research, audits)
Adversarial verificationFindings must be trusted before acting on themPer-finding refuter in Evaluate
Best-of-N or tournamentA single draft is unreliableCompeting Architect designs plus a judge
Planners → workers → judgeMany agents on one codebaseBatch runs across work items, one writer per module

The reliability kit for unattended runs is equally consistent across sources: a fresh context per unit of work, re-grounded from files; machine-readable progress; a commit per task; caps on turns and dollars plus a no-progress stop; deterministic replay; logs written for agents (Carlini: “put the reason on the same line so grep will find it”); isolation through worktrees and least-privilege credentials; and scheduled cleanup and docs-drift loops.

Two details matter for Faber’s executor. The Agent SDK’s maxBudgetUsd “counts only the call’s own spend; totals restored from a resumed session don’t count,” and no wall-clock limit appears among its options, so Faber needs its own cumulative run budget and per-step timeouts (Agent SDK TypeScript reference). And because workflows accept “no mid-run user input,” sign-off between stages means running “each stage as its own workflow,” which fits Faber’s phase gates if Faber ever compiles phases to Claude Code workflow scripts.

The broader implication is that Faber’s code-driven executor should be the authoritative control plane for unattended runs, with every step returning schema-validated output the code can branch on, while the single-session LLM orchestrator becomes an interactive mode or a planner that emits plan data for the executor to run.

Validators gate runs; only evals show whether the gates work

Per-step validators and evals do different jobs, and one cannot stand in for the other. Hamel Husain and Shreya Shankar distinguish guardrails, inline checks in the request/response path where false positives are treated as production bugs, from evaluators, which run after a response is produced, feed dashboards, regression tests and improvement loops, and do not block the original answer (Husain and Shankar, evals FAQ).

Faber’s LLM validators are a hybrid: in-loop judges that block. That makes each one a component whose accuracy must be measured, not a measurement of the system. A validator judges one run; it cannot say whether the pipeline improved after a prompt change, a model upgrade or a rewrite of the validator itself. And if the pipeline’s own validator also grades the eval suite, a lenient validator inflates both the gate and the score, lining up the holes in what Anthropic calls the Swiss cheese model, where “no single evaluation layer catches every issue” (Anthropic, January 2026). Offline evals need independent ground truth: hidden tests, reference solutions and human labels.

The program: small, real and repeated

Start with error analysis on about 100 real traces, stopping when “no new failure types” appear in the last 20 (Husain, error-analysis skill). Then, following Anthropic’s guidance (Anthropic, January 2026):

  • Build “20-50 simple tasks drawn from real failures,” each with a reference solution that passes every grader, and unambiguous enough that “two domain experts would independently reach the same pass/fail verdict.”
  • Run each trial from a clean environment, because Anthropic caught Claude “gaining an unfair advantage on some tasks by examining the git history from previous trials.”
  • Grade outcomes rather than tool-call paths. Use deterministic graders first and calibrated LLM judges only for judgment calls.
  • Split the suite into capability evals that “should start at a low pass rate” and regression evals that should sit near 100%.
  • Above all, read transcripts: “we do not take eval scores at face value until someone digs into the details of the eval and reads some transcripts.”

The graders themselves fail. A frontier model “initially scored 42% on CORE-Bench” until rigid grading, ambiguous task specs and irreproducible tasks were fixed; with a less constrained harness, it scored 95% (same source). OpenAI stopped treating SWE-bench Verified as a measure of frontier capability after an audit found material issues in test design or problem description in 59.4% of 138 hard tasks (OpenAI, February 2026).

Small suites detect only large effects

The arithmetic is worth internalizing. Anthropic justifies 20–50 tasks because early changes have a “large effect size.” By the standard two-proportion power calculation (5% significance, 80% power, pass rates near 50%), comparing two pipeline versions on independent task sets needs about 44 tasks per arm to detect a 30-point change, 98 for 20 points and 392 for 10 points; running both versions on the same tasks with a 0.5 correlation roughly halves those counts. With 50 tasks, the 95% confidence interval on a pass rate near 50% is about ±14 points. These are computed figures, a model rather than a measurement.

Anthropic’s statistics guidance adds paired comparisons and clustered standard errors, which “can be over three times as large as naive standard errors” (Anthropic, November 2024). Extra trials per task tighten each task’s estimate but do not replace more tasks. Faber’s spec proposes a pilot of 10 past work items × 3 runs before sizing the full suite; a suite of 15–20 tasks is the right size for catching regressions and gross wins, and too small to resolve a validator tweak worth ten points.

For unattended runs, the metric is pass^k, not pass@k

Anthropic defines pass@k as the chance that at least one of k attempts succeeds and pass^k as the chance that all k succeed: “If your agent has a 75% per-trial success rate and you run 3 trials, the probability of passing all three is (0.75)³ ≈ 42%,” so pass@k suits tools “where one success matters” and pass^k suits agents “where consistency is essential” (Anthropic, January 2026). No human picks the best of k unattended runs; each run is one draw, and a 90%-reliable agent run ten times has only about a 35% chance of zero failures.

Real benchmarks decay more slowly than p^k because difficulty varies by task. On tau-bench retail, one model fell from 0.692 at pass^1 to 0.462 at pass^4, and another’s 0.383 at pass^4 is far above the 0.133 its 0.604 pass^1 rate would predict if every task were equally hard (tau-bench). As k grows, pass^k converges on the share of tasks an agent handles every time, so Faber should report each task as always, sometimes or never solved, and target the “sometimes” tasks.

Phases compound: six phases at 95% each yield about 74% end-to-end per run and about 40% at pass^3, so reliability must be measured end to end, not phase by phase.

Validator accuracy drives delivered quality more than intuition suggests

Here is a model, not a measurement. Take a generator that succeeds 60% of the time, a validator that passes 90% of good outputs (its true-positive rate) and fails 80% of bad ones (its true-negative rate), with retries until acceptance: about 87% of accepted outputs are good, after 1.6 attempts on average. Raising the true-negative rate to 95% lifts that to 96% at 1.8 attempts. Across five gated phases, the difference compounds to roughly 50% versus 83% end to end.

The rate at which a validator catches bad work dominates delivered quality; the rate at which it passes good work mainly drives retries and cost.

Cost per successful task should therefore count generation, validation, retries and human review across all runs, divided by the runs that independent graders verify as correct. At $2.00 per attempt and $0.30 per validation, the first scenario costs about $3.71 per accepted output and $4.26 per output that is actually good. Princeton’s “AI Agents That Matter” argues that agent evaluations must be cost-controlled, noting that complex agent designs failed to beat simple baselines on HumanEval despite costing far more (Kapoor et al., 2024).

Validators need their own evaluation

Against human labels, Husain’s protocol splits labeled examples into train, dev and test sets, targets true-positive and true-negative rates above 90% each, forbids iterating after seeing test results and pins model versions (Husain, validate-evaluator skill).

Against planted defects, the CriticGPT method of planting subtle bugs in model-written code (OpenAI, CriticGPT) generalizes to dropped requirements, weakened tests and scope creep. Qodo’s vendor-run benchmark of injected issues found that “most tools cluster toward high precision and low recall” (Qodo benchmark), which is exactly the failure a gate cannot afford. Anthropic’s under-1% incorrect rate for Code Review measures precision only.

Judge construction follows the same evidence: narrow binary checks per dimension, reasoning before the verdict, a different model from the generator (“generally best practice,” per Anthropic’s own docs (Claude platform docs)), and a majority of three votes. One open conflict: Anthropic’s January 2026 guidance favors an isolated judge per dimension, while its June 2025 research system found a single judge call most consistent (Anthropic, June 2025), so Faber should test both on its own labeled set.

Does each piece still earn its cost?

Claude Code’s claude plugin eval runs every case with and without the plugin, three times by default, and reports the difference, because “if a case scores 1.0 both with and without the plugin, the plugin isn’t what made it pass.” Its docs also advise pinning the agent and judge models in CI “so a model rollout isn’t mistaken for a plugin regression” (Claude Code plugin evals). The same with/without design is the right test for each Faber validator, contract step and skill, re-run at every model upgrade.

Production signals close the loop: Claude Code exports edit accept/reject decisions and cost through OpenTelemetry (Claude Code monitoring), and DORA’s deployment rework rate offers an outcome measure for agent-shipped changes (DORA).

Nine changes for Faber, ranked by evidence

Faber’s harness hardening spec already proposes most of the mechanics: parsing validator verdicts, evidence-gated completion, a floor guard, fresh-context review with a capped fix loop, and an eval pilot. The research supports that sequencing and sharpens it in four places. It favors making the code-driven executor the default control plane for unattended runs. It says a small suite will detect only large effects. It treats validators as components that need their own measurement, catch rate first. And it calls for an explicit, frozen layering of criteria so that one validator can serve every project.

#ChangeEvidence behind itConfidence
1Make verdicts gate: parse a schema-validated verdict from every step, fail closed, and treat “could not run” and “unverified” as non-passesAnthropic’s per-criterion hard thresholds; the floor guard’s 0/1/2 exit contract; workflows reporting uncheckable claims as unverifiedHigh
2One fractary-faber verify --stage entrypoint run by the executor, hooks and CI, writing evidence files that completion requiresSpec Kit gates; agent-skills check tiers; Claude Code Stop/TaskCompleted hooksHigh
3Fresh-context validators that receive artifact, contract, evidence and tools, never the implementer’s claims, scoped to correctness and stated requirementsAnthropic’s harness work; Claude Code docs; doubt-driven development; self-preference and shared-context researchHigh on direction; effect size on Faber’s own work still to be measured
4Layered done: floor as a managed hook, human-ratified project profile, L2 criteria frozen at the Frame gate, L3 verify steps coverage-checked in Architect, optional Build contractagent-skills, Spec Kit, OpenSpec, Anthropic sprint contractsMedium-high; design consensus
5Bounded loops: at most three validator cycles, a no-progress stop, a fresh implementer after two failed fixes, best checkpoint kept, decomposition instead of higher capsDoubt-driven development; Claude Code docs; Anthropic iteration dataMedium; practitioner defaults
6Code-driven executor as the default for unattended runs; single-session mode for interactive work or as a planner emitting plan data; approvals only at stage boundariesClaude Code workflows; OpenAI Agents SDK; 12-Factor Agents; long-run case studiesMedium-high
7Eval suite of real tasks, at least three clean runs each, scored on end-to-end pass^k and cost per successful task; validators measured on labeled outputs and planted defects; with/without ablations at every model upgradeAnthropic eval guidance; Husain’s protocols; claude plugin evalHigh on need; small suites resolve only large effects
8Phase and validator knowledge in skills, thin generic agents, and an assumption register that records which model weakness each component compensates forAnthropic skills posts; harness ablation practice; skill-creator guidanceMedium
9A different-model-family reviewer for high-risk changes such as auth, migrations and irreversible release stepsSelf-preference and jury studies, at abstract levelMedium-low; adds cost

Cost should be budgeted, not feared. Multi-agent systems use about 15 times the tokens of chat (Anthropic, June 2025), but a scoped evaluator was about 8% of Anthropic’s harness spend, so the expensive part of adding judges is the rework they trigger, which is the point. Faber should track tokens and dollars per phase, reserve wide fan-out for phases where checks are cheap and the work is genuinely parallel (research, audits, review, mechanical migrations), and let each validator justify itself through the with/without comparison.

Conclusion

The center of gravity has moved from making agents generate better code to deciding, cheaply and credibly, that the code is done. The components that keep surviving model upgrades (layered criteria, deterministic gates, fresh-context judges with tools, and pass^k evals) all serve that decision, while components that coach the model (context resets, forced decomposition, long system prompts) keep getting deleted as models improve.

That is why the validator is the one piece of a harness unlikely to be erased by better models: as capability grows, teams push work to the new edge, which is exactly where Anthropic found the evaluator “continued to give real lift.” For a framework used across many projects, phases are not the differentiator, because every spec-driven tool has them. A portable, measured judge that reads a project-specific contract is.

That reframes Faber’s promise of “earned autonomy,” which its README describes as a progression over days and months. The evidence points to earning autonomy by measurement rather than elapsed time. A project or work type would move from assisted to autonomous when its end-to-end pass^3 and its validators’ measured catch rate clear agreed thresholds on its own eval slice, and move back when production rework rises.

The compounding arithmetic shows why calendar-based trust is unsafe: with a generator that succeeds 60% of the time per phase, validators that catch 80% of bad work leave only about half of five-phase runs free of an accepted defect. Faber’s eval suite, run with and without each component, is how I’ll measure what each of these practices is worth on my own work.


This is the eighth post in the Earned Autonomy series and the reference behind the others. Previous: pass^k: The Reliability Math of Unattended Agents. Start at the beginning: Who Decides an Agent’s Work Is Done? How Faber Is Changing.

The Stack

Build on the stack the studio builds with.

Open-source, Apache-2.0 tools for building agentic systems without vendor lock-in. The same stack behind every venture on this site.