Earned Autonomy · Part 8
Agent Harnesses and Evals: The Full Research Note
The full research note behind the Earned Autonomy series: how code owns the loop, a separate judge decides done, and evals show whether any of it helps.
Josh McWilliam
35 min read
- agentic-systems
- harness
- evals
- claude-code
- production
This is the research note behind the Earned Autonomy series. I wrote it while revamping Faber, the open-source workflow tool the studio builds with, to match what the field has learned about running agents unattended. Every claim links to its source.
Itâs long and dense on purpose. If you want the arguments one at a time, the series posts are the readable version:
- Who Decides an Agentâs Work Is Done? How Faber Is Changing
- The Agent That Did the Work Shouldnât Decide Itâs Done
- Done Is a Contract: Four Layers Your Agents Canât Rewrite
- What Running Thousands of Agents Actually Takes
- Your Agent Harness Is a List of What the Model Canât Do
- A Validator Is Not an Eval: How to Measure the Judge
- pass^k: The Reliability Math of Unattended Agents
The note reflects sources published up to early October 2026, and parts of it will date. The Faber changes it recommends are tracked in the public harness hardening spec.
Script the loop and separate the judge
Leading practice for autonomous software development has converged on one architecture. Deterministic code owns the outer loop: state, stage transitions, budgets, gates and human sign-off. The model owns judgment inside each step, and increasingly writes task-specific orchestration scripts that a runtime replays deterministically. And âdoneâ is decided by evidence and a separate judge rather than by the agent that did the work.
For Faber, four shifts follow.
- The definition of done should be a layered, file-based contract that one generic validator interprets, instead of validators rewritten per project: a framework floor, a human-ratified project profile, acceptance criteria drafted in Frame and frozen at the FrameâArchitect gate, and per-task verify commands written in Architect.
- Validators should run as fresh-context agents that receive the artifact, the contract and harness-captured evidence, but never the implementerâs reasoning or claims. Self-review is systematically lenient and shared context feeds reward hacking. Separation pays off most when the judge can execute checks itself.
- âDonât build agents, build skillsâ rejected bespoke per-domain agent architectures, not agent instances. Anthropic shipped workflows that run hundreds of subagents within months of that talk, and Boris Chernyâs âthousandsâ of nightly agents are machine-spawned runs fenced by cheap verification.
- Per-step validators are runtime gates, not evals. Only an offline suite of real tasks, run several times each in clean environments and scored on pass^k, can say whether a validator, prompt or model change helped, and whether the validators catch what they claim to catch.
Done is a four-layer contract that planning writes and code enforces
The tools that spell out completion most explicitly all separate a standing bar from per-item criteria. Addy Osmaniâs agent-skills puts it in one line: âA task is done only when its acceptance criteria are met and the standing Definition of Done is satisfied,â where acceptance criteria are âdefined when planning the taskâ and the Definition of Done is âdefined once for the projectâ and reused, because âa Definition of Done that is renegotiated every sprint is not a Definition of Doneâ (agent-skills, definition-of-done.md).
GitHub Spec Kit encodes the same stack as constitution â spec â plan â tasks, with a semver-versioned, ratified constitution (Spec Kit constitution command) and a plan-stage âConstitution Checkâ marked âGATE: Must pass before Phase 0 researchâ (Spec Kit plan template). OpenSpec uses project config.yaml â requirement specs â change deltas â tasks (OpenSpec customization). Anthropicâs long-running harnesses compress the lower layers into a JSON feature list written by an initializer agent (Anthropic, November 2025) and add a per-chunk âsprint contractâ that the generator and evaluator negotiate âbefore any code was writtenâ (Anthropic, March 2026). The evidence for the layering is the toolsâ own documentation, and the pattern holds across all of them.
| Layer | Holds | Drafted by | Ratified by | Faber home |
|---|---|---|---|---|
| L0 floor | Stack-agnostic rules: no new suppressions, skipped or deleted tests, stubs or secrets; the bar may not be weakened | Framework | Framework maintainers; shipped as a managed hook | Identical in every project |
| L1 project profile | Per dimension: command, threshold, direction, stage, applicability; exceptions with an owner and an expiry | Agent, from stack detection plus a short interview | A named human; versioned | Project config read by every phase |
| L2 acceptance criteria | Stable-ID behaviors (EARS, Given/When/Then, or SHALL plus scenarios) and an out-of-scope list | Frame agent | Human, or autonomy policy, at the FrameâArchitect gate | Frame output, hash-locked |
| L3 task verification | Per-task acceptance (three bullets at most) plus a verify command, mapped to L2 IDs | Architect agent | Deterministic coverage check | Architect output |
| L2.5 contract (optional) | Granular testable behaviors for one chunk of Build | Implementer proposes | Validator approves; may only tighten L0âL2 | Start of Build |
Planning phases matter because they are where criteria come from, but they should stop at the right altitude. Anthropic kept its planner through every simplification because âwithout the planner, the generator under-scoped,â yet confined it to product-level specs because âif the planner tried to specify granular technical details upfront and got something wrong, the errors in the spec would cascade into the downstream implementationâ (Anthropic, March 2026).
The sprint contract bridged that gap: the generator âproposed what it would build and how success would be verified, and the evaluator reviewed that proposal,â and one sprint alone carried 27 criteria, granular enough that the evaluatorâs failures pointed at specific handlers and routes.
For Faber, Frame should own what must be true (L2, testable and ID-stamped). Architect should own how each criterion will be proven (L3, with a coverage map back to L2 in the style of Spec Kitâs strictly read-only /analyze, which flags uncovered requirements that block baseline functionality as critical (Spec Kit analyze)). Build can open with a contract step for work at the edge of model capability. Anthropic dropped the sprint construct when a newer model arrived, so the contract step is model-dependent scaffolding; the layers themselves are not.
The weak point: the implementer still marks its own completion
Spec Kitâs implementer ticks [X] in its own task file, and its checklist gate can be waved through with a âyesâ (Spec Kit implement). OpenSpecâs verify step âdoes not block archive, but surfaces issuesâ (OpenSpec commands). Only an independent evaluatorâs verdict or a deterministic guard takes the decision away from the worker.
The mechanisms that stop criteria being quietly weakened are concrete:
- Keep status in JSON, because âthe model is less likely to inappropriately change or overwrite JSON files compared to Markdown filesâ (Anthropic, November 2025).
- Make post-build gap-filling append-only, as Spec Kitâs
/convergedoes (âAPPEND-ONLY, NEVER REWRITEâ) (Spec Kit converge). - Diff the bar against the branch point, as agent-skillsâ floor guard does, because âagents donât craft clever loopholes. They hit a red check and take the cheapest road to green,â so âtightening the bar should be silent; loosening it should be loudâ (agent-skills CDD skill).
The same skill ranks checks by circularity, from external tools such as axe or osv-scanner (âthe agent canât argue with theseâ) down to the projectâs own test suite (âthe only genuinely circular oneâ), and requires at least one external check per project.
Faber should hash-lock L1 and L2 at their approval gates, fail Evaluate closed if either changed without re-approval, give Build write access only to a status-and-evidence file, and route any change to the criteria back to Frame.
Validators stay generic when the project-specific part is data
The converging recipe is a single engine plus a per-project declaration in which every dimension names its command, threshold, direction (minimum, maximum or ratchet), stage and applicability. agent-skills insists that âevery row names the command that produces the verdict,â because âa dimension with a number and no command ⊠is an aspiration, not a constraintâ (agent-skills CDD skill).
The engine detects the stack before asking questions, replaces invented targets with ratchets (âSet 80% coverage on a codebase at 62% and you get a red build foreverâ), marks inapplicable dimensions explicitly (Lighthouse and axe âneed a URL,â so a CLI or library drops them rather than faking them), and scopes exceptions by path with an owner and an expiry. Customization flows through layered overrides, such as Spec Kitâs priority-ordered preset stack (Spec Kit presets) and Claude Codeâs hook scopes, not through forked validators.
In Faber terms, the validator becomes an interpreter that dispatches on each criterionâs verification method (test, command, end-to-end, HTTP, snapshot, policy, evaluator or manual) and reads commands from the project profile. Web, library and CLI presets follow directly from the sources; infrastructure-as-code and data-pipeline presets are reasonable extrapolations that none of the surveyed tools ships yet.
Enforcement has to live below the prompt
The robust pattern is one verify entrypoint run at every boundary. Spec Kitâs community gates extension promises âone policy file, one verify entrypoint, identical results at every boundaryâ across agent hooks, git hooks and CI (Spec Kit community catalog), and agent-skills splits the same checks into fast, task and full tiers, warning that âa check that stalls the agent gets switched off.â
Claude Codeâs Stop, SubagentStop and TaskCompleted hooks can refuse completion with exit code 2, but the harness overrides a stop hook after eight consecutive forced continuations, so a gate that cannot be satisfied must escalate rather than loop. Managed-policy settings with allowManagedHooksOnly let an organization ship a floor that projects cannot silently disable (Claude Code hooks).
A fractary-faber verify --stage command with a 0/1/2 exit contract, where 2 means âcould not runâ and never counts as a pass (âNever let a 2 read as a 0â) (agent-skills floor guard), wired into hooks, the executor and CI, would give Faber the same property.
Self-review fails exactly where validators matter most
Anthropicâs own harness work states the problem plainly: âWhen asked to evaluate work theyâve produced, agents tend to respond by confidently praising the workâeven when, to a human observer, the quality is obviously mediocre.â The fix was structural, not rhetorical: âSeparating the agent doing the work from the agent judging it proves to be a strong lever,â because âtuning a standalone evaluator to be skeptical turns out to be far more tractable than making a generator critical of its own work, and once that external feedback exists, the generator has something concrete to iterate againstâ (Anthropic, March 2026).
Claude Codeâs documentation gives the incentive version of the same argument: âClaude stops when the work looks done. Without a check it can run, âlooks doneâ is the only signal available,â and âa fresh context improves code review since Claude wonât be biased toward code it just wroteâ (Claude Code best practices).
The research literature explains why, through four reinforcing mechanisms. (The paper figures below are from abstracts unless marked as from the full text.)
Models are worse at finding their own errors than at fixing located ones. Intrinsic self-correction without external feedback struggles and âat timesâ degrades results (Huang et al., ICLR 2024). A TACL survey found self-correction works well in tasks that can use reliable external feedback, while no prior work showed it succeeding with feedback from prompted LLMs outside tasks exceptionally suited to it (Kamoi et al., 2024). Models backtrack successfully once told where the mistake is (Tyen et al., 2024). Code self-repair is âbottlenecked by the modelâs ability to provide feedback on its own code,â while a stronger modelâs feedback helps far more (Olausson et al., ICLR 2024). And models fix errors in user input but miss identical errors in their own output, a 64.5% average âblind spotâ across 14 open models (Self-Correction Bench, 2025).
Judges prefer what they recognize. One frontier model was 73.5% accurate at telling its own outputs from those of two other LLMs and humans (full text), and self-preference rises linearly with self-recognition (Panickssery et al., NeurIPS 2024), while low-perplexity, familiar text gets higher scores than human evaluators give it, whoever wrote it (Wataoka et al., 2024). Crucially, harmful self-preference persists in the cases where the evaluator erred as the generator, and stronger models show more of it when they err (Chen et al., 2025), so self-review is weakest precisely on the defects a validator exists to catch.
Shared context breeds reward hacking. When one model generated and evaluated in the same context window, evaluator ratings rose while human-judged quality stagnated or fell, and severity depended on how much context the two roles shared (Pan et al., 2024).
The agent that wants to stop is the one deciding it is done. Anthropic watched its evaluator âidentify legitimate issues, then talk itself into deciding they werenât a big deal and approve the work anyway.â
Separation is necessary but not sufficient
âOut of the box, Claude is a poor QA agent,â and Anthropicâs separate evaluator also tested superficially until the author spent âseveral roundsâ reading its logs and correcting its prompt against his own judgment (Anthropic, March 2026).
The large measured gains come from judges with an external signal. An agentic judge that inspects the workspace reached roughly 90% alignment with the human judgesâ consensus versus about 70% for a plain LLM judge on a 55-task development benchmark (full text) (Agent-as-a-Judge, 2024), and Anthropicâs evaluator drove the running app through Playwright and returned file:line causes rather than opinions.
Anthropicâs post names a single model for the whole harness, so its separation was of context, prompt and tools rather than model. A fresh context removes the shared-context factor; only a different model family or a tool-grounded check addresses familiarity bias, which is why panels of judges from disjoint model families beat a single large judge at over seven times lower cost (Verga et al., 2024). The case for separation rests on evidence converging from several directions, which is strong enough to act on.
Where Faber started
Faberâs phases, approvals and issue-to-pull-request flow were already in place. The hardening spec records where the work on âdoneâ begins: in plugin mode, every step, validators included, ran in one shared context, and in the CLI runtime, which already gives each step a fresh agent session, a stepâs verdict didnât yet gate the run (issue #235). That fix, verdicts that stop a run when they fail, has since merged and goes out in the next release (PR #243), along with saved run state and resume (PR #241). Moving validators into fresh contexts only helps once their verdicts gate, so parsing a structured verdict and failing closed comes first.
What a validator should see, and what it must never see
The sources agree on a four-part input packet and one exclusion. The validator gets:
- the artifact: diff, changed paths, commit, and how to build or run it;
- the contract: L1âL3 criteria, the out-of-scope list, project rules and severity definitions;
- evidence the harness captured itself: raw test logs, exit codes, screenshots;
- tools to gather its own: read-only repository access, a test runner, a browser.
It does not get the claim. Osmaniâs doubt-driven-development skill is blunt: âA fresh-context reviewer needs the artifact and the contract, not the journey⊠Strip your reasoning. If you hand over conclusions, youâll get back validation of your conclusionsâ (agent-skills doubt-driven-development). Claude Codeâs docs describe the same boundary: a fresh-context reviewer âsees only the diff and the criteria you give it, not the reasoning that produced the change, so it evaluates the result on its own termsâ (Claude Code best practices).
Anthropicâs official code-review plugin does pass each reviewer the pull requestâs title and description, to âprovide context regarding the authorâs intentâ (claude-code code-review command). The two positions reconcile if Faber passes intent but scrubs assertions such as âall tests passâ or âI verified Xâ from the handoff. The rule of thumb: give âwhat should be trueâ and withhold âwhy the author believes it is true.â
Evidence must be produced, not reported. Claude Codeâs /goal evaluator illustrates the limit of transcript-only judging: it âdoesnât run commands or read files independently,â so it can only be as good as what the worker chose to surface (Claude Code /goal). A gating validator should either run the checks itself or read evidence files the harness wrote, following the example verification rule in Anthropicâs managed Code Review docs, âbehavior claims need a file:line citation in the source, not an inference from namingâ (Claude Code Code Review).
Missing context is a common source of false positives, and the sources offer two designs: a diff-only reviewer told âdo not flag issues that you cannot validate without looking at context outside of the git diffâ (claude-code code-review command), which is cheaper with fewer findings, or full-codebase access plus a per-finding verification pass, which costs more and catches more (Claude Code Code Review). Faber can pick by risk tier.
Output should be per-criterion verdicts with severity, an explicit âUnknownâ escape hatch that Anthropic recommends for judges (Anthropic, January 2026), and an âunverifiedâ state that is never counted as a pass, mirroring how Claude Code workflows report claims their verifiers could not check (Claude Code workflows). Scope matters too: âA reviewer prompted to find gaps will usually report some, even when the work is sound,â and chasing every finding produces over-engineering, so validators should flag only gaps against correctness and the stated requirements (Claude Code best practices).
Two to three review cycles, then escalate or decompose
Practitioner bounds cluster at two to three LLM review cycles per artifact. Doubt-driven development stops after three (âescalate to user, donât grind a fourth aloneâ) and treats a need for more as a sizing failure: âthe artifact is too big â return to Step 2 and decompose. Do not lift the bound.â It also names âdoubt theater,â where repeated cycles surface findings that are never acted on, as a signal to stop (agent-skills doubt-driven-development).
Claude Code advises clearing context âafter two failed correctionsâ and re-prompting with what was learned (Claude Code best practices), and its managed reviewer suggests suppressing new nits after the first review so a one-line fix does not reach âround seven on style alone.â Anthropicâs subjective-design loop ran 5 to 15 iterations but plateaued, and the author âregularly saw cases where I preferred a middle iteration over the last one,â which argues for keeping the best-scoring checkpoint rather than the last (Anthropic, March 2026).
Faberâs spec caps fix cycles at three, then escalates or splits the work, which sits inside this band. Adding a no-progress stop (same finding recurring, unchanged artifact) and a fresh implementer context after two failed fixes would close the remaining loopholes.
The evaluator is cheap relative to what it protects
In Anthropicâs simplified harness, the three QA rounds cost $10.39 of a $124.70 run (about 8%) and 25 of roughly 230 minutes, while the build rounds dominated spend. QA still caught that âaudio recording is still stub-onlyâ and that core features were âdisplay-onlyâ (Anthropic, March 2026). Anthropicâs conclusion is conditional: the evaluator âis worth the cost when the task sits beyond what the current model does reliably solo.â
Its managed PR reviewer costs $15â25 per review and raised the share of PRs receiving substantive comments from 16% to 54%, with under 1% of findings marked incorrect by engineers, a vendor-reported precision figure without a recall measurement (Claude blog, Code Review).
Fresh contexts pay off for reads, verification and parallel transforms, not coupled writes
Validators are not the only steps that benefit from isolation. Work belongs in a separate context when it is read-heavy (exploration, search, audit); when it produces verbose output that only needs a summary (subagents typically return a âcondensed, distilled summary⊠often 1,000-2,000 tokensâ (Anthropic, September 2025)); when it is independent verification; when it is a per-item transform with an objective per-item check (file-disjoint, in worktrees); or when it needs narrower permissions, such as a read-only reviewer.
Work belongs in one context, or with a single writer, when decisions are tightly coupled, when reasoning is sequential, or when phases need constant back-and-forth. Cognition, whose June 2025 âDonât Build Multi-Agentsâ was the fieldâs main counterweight, revisited it in April 2026: âmulti-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions. A clean-context reviewer catches bugs the coder canât seeâ (Cognition, April 2026). A study of 260 agent configurations across six benchmarks found multi-agent performance ranging âfrom +80.8% on decomposable financial reasoning to -70.0% on sequential planning,â and that âarchitectures without centralized verification tend to propagate errors more than those with centralized coordinationâ (Kim et al., December 2025). Anthropicâs own multi-agent research system beat a single agent by 90.2% but used about 15Ă the tokens of chat, and was explicitly a poor fit for âmost coding tasks,â which have fewer truly parallelizable parts (Anthropic, June 2025).
Mapped onto Faber, this gives each phase a shape:
- Frame fans out read-only exploration and has one synthesizer write the acceptance criteria.
- Architect can generate competing designs in parallel, but one decider records the decision in an artifact every downstream step receives in full, honoring Cognitionâs rule to share context rather than fragment it.
- Build keeps one writer per module, parallelizing only across partitioned, independently testable units in separate worktrees, with an integration gate on the full suite.
- Evaluate runs fresh-context validators and, for high-stakes findings, adversarial refuters.
Across all phases, a fresh session per step suits long unattended runs, because subagents and fresh sessions inherit no conversation history (Claude Code subagents). The phase artifacts, not the chat transcript, must carry everything the next step needs.
Skills, agents and workflows are layers, not rivals
âDonât build agents, build skillsâ was an argument about architecture, not about how many agents to run. Barry Zhang and Mahesh Murag made it at the AI Engineer Code conference in November 2025 (AI Engineer talk page), and Anthropicâs January 22, 2026 write-up spells out what was abandoned: âWe used to think agents in different domains would look very different. A coding agent, a research agent, one for finance, one for marketingâeach seemed to need its own tools and scaffolding.â Instead, âClaude Code is a coding agent, but also a general-purpose agent that happens to work through code,â and skills âturn a capable generalist into a knowledgeable specialist.â
The postâs summary of the stack is the clearest statement of the layering: âthe loop reasons, the runtime executes, MCP connects, and skills guideâ (Claude blog, January 2026). What the industry moved away from is the bespoke agent class per domain, each with hand-built scaffolding. What it kept, and multiplied, is agent instances running one general harness.
Anthropicâs documentation treats the pieces as composable layers. Skills hold procedural knowledge, subagents handle task delegation and MCP handles tool connectivity, and âa code-review subagent can use Skills for language-specific best practicesâ (Claude blog, Skills explained). In Claude Code a subagent can preload skills, a skill marked context: fork runs inside a fresh subagent (Claude Code skills), and the bundled /batch command is itself a skill that splits a change across 5 to 30 worktree-isolated subagents (Claude Code, run agents in parallel). Managed Agents defines an agent as âthe model, system prompt, tools, MCP servers, and skillsâ (Managed Agents overview).
Barry Zhang, credited on the January skills post, also co-wrote the multi-agent research post, and the company launched dynamic workflows on May 28, 2026, in which âClaude dynamically writes orchestration scripts that run tens to hundreds of parallel subagents in a single session, checking its work before anything reaches youâ (Claude blog, May 2026). The docs now sort the primitives by who holds the plan: Claude turn by turn (subagents and skills), a lead agent (agent teams), or a script (workflows) (Claude Code workflows). A companion page chooses among them by three questions: who coordinates, whether workers need to talk, and whether they touch the same files, and still marks agent teams experimental (Claude Code, run agents in parallel).
Chernyâs thousands are machine-spawned runs fenced by verification
Boris Chernyâs public numbers climbed over a few months, and most of them reach us through press coverage of talks rather than transcripts.
- In January, coverage of his X thread described 5 terminal sessions plus 5â10 web sessions, and his advice to give Claude a way to verify its work, which he argued improves final quality â2-3xâ (VentureBeat, January 2026). That multiplier is an informal estimate. Anthropicâs current docs keep the advice and drop the number: a check Claude can run is âthe difference between a session you watch and one you walk away fromâ (Claude Code best practices).
- In a Sequoia Capital interview in May he still ran âfive to 10 sessions,â but âusually, every night, I have like a few thousand that are doing kind of deeper workâ (Business Insider, May 2026).
- In June, The New Stack reported that he no longer prompts Claude directly and that âmy job is to write loopsâ (The New Stack, June 2026).
- At Fortune Brainstorm Tech on June 8: âThis morning I was managing maybe a few hundred⊠Some days itâs ⊠thousands, or tens of thousands,â and he had not written code by hand in about eight months (Fortune, June 2026).
- At Metaâs @Scale he described a transition âto the point where agents are prompting agents that then write the code,â called loops âjust as important and as big a stepâ as the move from hand-written code to agents, and gave examples of standing loops that hunt for architecture improvements and duplicated abstractions (TechCrunch, June 2026).
A widely shared line, âyou build the harness that runs the loops,â has no primary source; it circulated through a third-party post and reads as a paraphrase of âmy job is to write loops.â
Read together, the numbers say less about human attention than about automation. The sessions Cherny supervises stayed at roughly 5â10; the hundreds and thousands are subagent runs inside workflows, loops and routines. Every large-scale example on record rests on a cheap, objective oracle.
- Nicholas Carliniâs 16 parallel agents ran about 2,000 sessions over two weeks for just under $20,000 to produce a 100,000-line C compiler that builds Linux 6.9, with GCC as a differential oracle and the warning that âitâs important that the task verifier is nearly perfectâ (Anthropic, February 2026).
- Bunâs Zig-to-Rust rewrite ran 64 agents (four workflow shards of 16), with the compilerâs error output as the work queue and the test suite as the bar (The Pragmatic Engineer, July 2026).
- Workflows default to 16 concurrent agents, configurable up to 256, with a hard cap of 1,000 per run (Claude Code workflows).
Scale is the product of scripted orchestration and machine verification, not of more supervision.
For Faber, the lesson is to put knowledge in skills and keep agents generic. Phase procedures, rubrics, the done-criteria interpreterâs instructions and each validatorâs criteria belong in versioned, progressively disclosed skills that can be preloaded into a validator subagent. Agent definitions should stay thin and exist mainly to scope permissions, such as a read-only reviewer. A validator then becomes âgeneric subagent + validator skill + fresh context + schema-validated output,â not a hand-built agent class per phase.
Anthropicâs skill-creator guidance adds a durability test: âcapability upliftâ skills become obsolete once the base model âstarts passing your evals without the skill loaded,â while âencoded preferenceâ skills stay useful (Claude blog, skill-creator). Faberâs project-specific done profiles are the durable kind; its generic âhow to code wellâ prompts are the perishable kind and should be re-tested at every model upgrade.
Harnesses are thinning while verification and infrastructure thicken
The control-flow question is largely settled, with one twist. Claude Codeâs workflow documentation states the consensus: âA workflow moves the plan into code. With subagents, skills, and agent teams, Claude is the orchestrator: it decides turn by turn what to spawn or assign next, and every result goes into a context window. A workflow script holds the loop, the branching, and the intermediate results itself, so Claudeâs context holds only the final answer.â The runtime even makes Date.now() and Math.random() throw inside scripts âso that a relaunched run repeats the same agent() callsâ (Claude Code workflows).
OpenAIâs Agents SDK docs describe code-based orchestration as more deterministic and predictable than LLM-based routing (OpenAI Agents SDK). HumanLayerâs 12-Factor Agents makes âown your control flowâ a principle so that loops can pause for humans and resume durably (12-Factor Agents). And the long-run case studies, from Geoffrey Huntleyâs Ralph loop to Anthropicâs initializer harness and Carliniâs compiler fleet, keep the outer loop in bash or code and state in files and git (how-to-ralph-wiggum).
The counter-evidence is narrower than it looks. LLM lead agents work for breadth-first research; Cursor found that planners, workers and a judge agent outperformed flat self-coordination at fleet scale (Cursor, January 2026); and Anthropicâs April 2026 guidance tells builders to âlet Claude orchestrate its own actionsâ through code execution rather than hand-written tool plumbing (Claude blog, April 2026).
The twist is that the model now writes the orchestration: the model writes the graph, and a deterministic runtime executes and replays it. âGraph engineeringâ is not an established term; the field says workflows versus agents, code versus LLM orchestration, dynamic workflows, harness engineering (popularized by OpenAIâs February 2026 post and Birgitta Böckelerâs articles) and, most recently, loop engineering.
Components that compensate for model weaknesses keep going stale
Anthropicâs principle is that âevery component in a harness encodes an assumption about what the model canât do on its own, and those assumptions are worth stress testing⊠because they can quickly go stale as models improve.â Its own harness dropped context resets once a newer model stopped wrapping up early near its context limit, then dropped sprint decomposition with the generation after that and moved evaluation to a single end-of-run pass. A radical cut failed, so the method that worked was removing one component at a time (Anthropic, March 2026).
The conclusion was not that harnesses vanish: âthe space of interesting harness combinations doesnât shrink as models improve. Instead, it moves.â Minimal agents make the same point from the other side: mini-swe-agent is about 100 lines of Python using only bash and scores above 74% on SWE-bench Verified (mini-swe-agent).
What does not go stale is governance and infrastructure: oracles and verification, security boundaries, durable state and resume, budget caps, isolation and approval gates. Böckelerâs split of a harness into feedforward âguidesâ and feedback âsensorsâ adds a sobering limit: neither deterministic nor LLM-based sensors reliably catch âmisdiagnosis of issues, overengineering and unnecessary features, misunderstood instructionsâ (Böckeler, martinfowler.com), which is the strongest argument for keeping a human at the framing gate and at outcome review.
Five directions
- Harnesses are becoming per-task and model-authored. With ultracode, âa single request can turn into several workflows in a row: one to understand the code, one to make the change, and one to verify itâ (Claude Code workflows).
- Execution is moving onto managed runtimes that separate the âbrain,â the sandboxed âhandsâ and a durable session log outside the context window, with credentials ânever reachable from the sandbox where Claudeâs generated code runsâ (Anthropic, Managed Agents, April 2026). Managed Agents also offers âoutcomes,â which define done through a rubric and a grader that runs in a separate context window, a hosted version of the validator pattern (Managed Agents, define outcomes).
- Agents are becoming event-driven. Routines fire on schedules, API calls or GitHub events; their fired prompts âcanât act as approval or consent,â and âa green status⊠does not mean the task in your prompt succeededâ (Claude Code routines).
- Effort is shifting to the environment. OpenAIâs harness-engineering account describes about a million lines and roughly 1,500 PRs with zero hand-written code, a short AGENTS.md acting as a table of contents, and custom linters whose error messages carry remediation instructions for agents (OpenAI, February 2026).
- Verification is the bottleneck. Carlini warns that âthe thought of programmers deploying software theyâve never personally verified is a real concernâ (Anthropic, February 2026), and the Agent SDK guidance ranks rule-based feedback first while calling LLM-as-judge âgenerally not a very robust methodâ (Claude blog, September 2025).
The named patterns, as step types
| Pattern | Use it when | Faber step type |
|---|---|---|
| Loop until a check passes (evaluator-optimizer) | Objective criteria exist: type check, tests, rubric | Verifyâfix loop with a cap and a âtwo rounds without progressâ stop |
| Fresh-session task loop (Ralph, initializer/coder) | Long builds with backpressure from tests, types and build | Build as one task per session, JSON task state, one commit per task |
| Fan-out and synthesize | Many independent items or sources | Parallel read-only step with one synthesizer (Frame research, audits) |
| Adversarial verification | Findings must be trusted before acting on them | Per-finding refuter in Evaluate |
| Best-of-N or tournament | A single draft is unreliable | Competing Architect designs plus a judge |
| Planners â workers â judge | Many agents on one codebase | Batch runs across work items, one writer per module |
The reliability kit for unattended runs is equally consistent across sources: a fresh context per unit of work, re-grounded from files; machine-readable progress; a commit per task; caps on turns and dollars plus a no-progress stop; deterministic replay; logs written for agents (Carlini: âput the reason on the same line so grep will find itâ); isolation through worktrees and least-privilege credentials; and scheduled cleanup and docs-drift loops.
Two details matter for Faberâs executor. The Agent SDKâs maxBudgetUsd âcounts only the callâs own spend; totals restored from a resumed session donât count,â and no wall-clock limit appears among its options, so Faber needs its own cumulative run budget and per-step timeouts (Agent SDK TypeScript reference). And because workflows accept âno mid-run user input,â sign-off between stages means running âeach stage as its own workflow,â which fits Faberâs phase gates if Faber ever compiles phases to Claude Code workflow scripts.
The broader implication is that Faberâs code-driven executor should be the authoritative control plane for unattended runs, with every step returning schema-validated output the code can branch on, while the single-session LLM orchestrator becomes an interactive mode or a planner that emits plan data for the executor to run.
Validators gate runs; only evals show whether the gates work
Per-step validators and evals do different jobs, and one cannot stand in for the other. Hamel Husain and Shreya Shankar distinguish guardrails, inline checks in the request/response path where false positives are treated as production bugs, from evaluators, which run after a response is produced, feed dashboards, regression tests and improvement loops, and do not block the original answer (Husain and Shankar, evals FAQ).
Faberâs LLM validators are a hybrid: in-loop judges that block. That makes each one a component whose accuracy must be measured, not a measurement of the system. A validator judges one run; it cannot say whether the pipeline improved after a prompt change, a model upgrade or a rewrite of the validator itself. And if the pipelineâs own validator also grades the eval suite, a lenient validator inflates both the gate and the score, lining up the holes in what Anthropic calls the Swiss cheese model, where âno single evaluation layer catches every issueâ (Anthropic, January 2026). Offline evals need independent ground truth: hidden tests, reference solutions and human labels.
The program: small, real and repeated
Start with error analysis on about 100 real traces, stopping when âno new failure typesâ appear in the last 20 (Husain, error-analysis skill). Then, following Anthropicâs guidance (Anthropic, January 2026):
- Build â20-50 simple tasks drawn from real failures,â each with a reference solution that passes every grader, and unambiguous enough that âtwo domain experts would independently reach the same pass/fail verdict.â
- Run each trial from a clean environment, because Anthropic caught Claude âgaining an unfair advantage on some tasks by examining the git history from previous trials.â
- Grade outcomes rather than tool-call paths. Use deterministic graders first and calibrated LLM judges only for judgment calls.
- Split the suite into capability evals that âshould start at a low pass rateâ and regression evals that should sit near 100%.
- Above all, read transcripts: âwe do not take eval scores at face value until someone digs into the details of the eval and reads some transcripts.â
The graders themselves fail. A frontier model âinitially scored 42% on CORE-Benchâ until rigid grading, ambiguous task specs and irreproducible tasks were fixed; with a less constrained harness, it scored 95% (same source). OpenAI stopped treating SWE-bench Verified as a measure of frontier capability after an audit found material issues in test design or problem description in 59.4% of 138 hard tasks (OpenAI, February 2026).
Small suites detect only large effects
The arithmetic is worth internalizing. Anthropic justifies 20â50 tasks because early changes have a âlarge effect size.â By the standard two-proportion power calculation (5% significance, 80% power, pass rates near 50%), comparing two pipeline versions on independent task sets needs about 44 tasks per arm to detect a 30-point change, 98 for 20 points and 392 for 10 points; running both versions on the same tasks with a 0.5 correlation roughly halves those counts. With 50 tasks, the 95% confidence interval on a pass rate near 50% is about ±14 points. These are computed figures, a model rather than a measurement.
Anthropicâs statistics guidance adds paired comparisons and clustered standard errors, which âcan be over three times as large as naive standard errorsâ (Anthropic, November 2024). Extra trials per task tighten each taskâs estimate but do not replace more tasks. Faberâs spec proposes a pilot of 10 past work items Ă 3 runs before sizing the full suite; a suite of 15â20 tasks is the right size for catching regressions and gross wins, and too small to resolve a validator tweak worth ten points.
For unattended runs, the metric is pass^k, not pass@k
Anthropic defines pass@k as the chance that at least one of k attempts succeeds and pass^k as the chance that all k succeed: âIf your agent has a 75% per-trial success rate and you run 3 trials, the probability of passing all three is (0.75)Âł â 42%,â so pass@k suits tools âwhere one success mattersâ and pass^k suits agents âwhere consistency is essentialâ (Anthropic, January 2026). No human picks the best of k unattended runs; each run is one draw, and a 90%-reliable agent run ten times has only about a 35% chance of zero failures.
Real benchmarks decay more slowly than p^k because difficulty varies by task. On tau-bench retail, one model fell from 0.692 at pass^1 to 0.462 at pass^4, and anotherâs 0.383 at pass^4 is far above the 0.133 its 0.604 pass^1 rate would predict if every task were equally hard (tau-bench). As k grows, pass^k converges on the share of tasks an agent handles every time, so Faber should report each task as always, sometimes or never solved, and target the âsometimesâ tasks.
Phases compound: six phases at 95% each yield about 74% end-to-end per run and about 40% at pass^3, so reliability must be measured end to end, not phase by phase.
Validator accuracy drives delivered quality more than intuition suggests
Here is a model, not a measurement. Take a generator that succeeds 60% of the time, a validator that passes 90% of good outputs (its true-positive rate) and fails 80% of bad ones (its true-negative rate), with retries until acceptance: about 87% of accepted outputs are good, after 1.6 attempts on average. Raising the true-negative rate to 95% lifts that to 96% at 1.8 attempts. Across five gated phases, the difference compounds to roughly 50% versus 83% end to end.
The rate at which a validator catches bad work dominates delivered quality; the rate at which it passes good work mainly drives retries and cost.
Cost per successful task should therefore count generation, validation, retries and human review across all runs, divided by the runs that independent graders verify as correct. At $2.00 per attempt and $0.30 per validation, the first scenario costs about $3.71 per accepted output and $4.26 per output that is actually good. Princetonâs âAI Agents That Matterâ argues that agent evaluations must be cost-controlled, noting that complex agent designs failed to beat simple baselines on HumanEval despite costing far more (Kapoor et al., 2024).
Validators need their own evaluation
Against human labels, Husainâs protocol splits labeled examples into train, dev and test sets, targets true-positive and true-negative rates above 90% each, forbids iterating after seeing test results and pins model versions (Husain, validate-evaluator skill).
Against planted defects, the CriticGPT method of planting subtle bugs in model-written code (OpenAI, CriticGPT) generalizes to dropped requirements, weakened tests and scope creep. Qodoâs vendor-run benchmark of injected issues found that âmost tools cluster toward high precision and low recallâ (Qodo benchmark), which is exactly the failure a gate cannot afford. Anthropicâs under-1% incorrect rate for Code Review measures precision only.
Judge construction follows the same evidence: narrow binary checks per dimension, reasoning before the verdict, a different model from the generator (âgenerally best practice,â per Anthropicâs own docs (Claude platform docs)), and a majority of three votes. One open conflict: Anthropicâs January 2026 guidance favors an isolated judge per dimension, while its June 2025 research system found a single judge call most consistent (Anthropic, June 2025), so Faber should test both on its own labeled set.
Does each piece still earn its cost?
Claude Codeâs claude plugin eval runs every case with and without the plugin, three times by default, and reports the difference, because âif a case scores 1.0 both with and without the plugin, the plugin isnât what made it pass.â Its docs also advise pinning the agent and judge models in CI âso a model rollout isnât mistaken for a plugin regressionâ (Claude Code plugin evals). The same with/without design is the right test for each Faber validator, contract step and skill, re-run at every model upgrade.
Production signals close the loop: Claude Code exports edit accept/reject decisions and cost through OpenTelemetry (Claude Code monitoring), and DORAâs deployment rework rate offers an outcome measure for agent-shipped changes (DORA).
Nine changes for Faber, ranked by evidence
Faberâs harness hardening spec already proposes most of the mechanics: parsing validator verdicts, evidence-gated completion, a floor guard, fresh-context review with a capped fix loop, and an eval pilot. The research supports that sequencing and sharpens it in four places. It favors making the code-driven executor the default control plane for unattended runs. It says a small suite will detect only large effects. It treats validators as components that need their own measurement, catch rate first. And it calls for an explicit, frozen layering of criteria so that one validator can serve every project.
| # | Change | Evidence behind it | Confidence |
|---|---|---|---|
| 1 | Make verdicts gate: parse a schema-validated verdict from every step, fail closed, and treat âcould not runâ and âunverifiedâ as non-passes | Anthropicâs per-criterion hard thresholds; the floor guardâs 0/1/2 exit contract; workflows reporting uncheckable claims as unverified | High |
| 2 | One fractary-faber verify --stage entrypoint run by the executor, hooks and CI, writing evidence files that completion requires | Spec Kit gates; agent-skills check tiers; Claude Code Stop/TaskCompleted hooks | High |
| 3 | Fresh-context validators that receive artifact, contract, evidence and tools, never the implementerâs claims, scoped to correctness and stated requirements | Anthropicâs harness work; Claude Code docs; doubt-driven development; self-preference and shared-context research | High on direction; effect size on Faberâs own work still to be measured |
| 4 | Layered done: floor as a managed hook, human-ratified project profile, L2 criteria frozen at the Frame gate, L3 verify steps coverage-checked in Architect, optional Build contract | agent-skills, Spec Kit, OpenSpec, Anthropic sprint contracts | Medium-high; design consensus |
| 5 | Bounded loops: at most three validator cycles, a no-progress stop, a fresh implementer after two failed fixes, best checkpoint kept, decomposition instead of higher caps | Doubt-driven development; Claude Code docs; Anthropic iteration data | Medium; practitioner defaults |
| 6 | Code-driven executor as the default for unattended runs; single-session mode for interactive work or as a planner emitting plan data; approvals only at stage boundaries | Claude Code workflows; OpenAI Agents SDK; 12-Factor Agents; long-run case studies | Medium-high |
| 7 | Eval suite of real tasks, at least three clean runs each, scored on end-to-end pass^k and cost per successful task; validators measured on labeled outputs and planted defects; with/without ablations at every model upgrade | Anthropic eval guidance; Husainâs protocols; claude plugin eval | High on need; small suites resolve only large effects |
| 8 | Phase and validator knowledge in skills, thin generic agents, and an assumption register that records which model weakness each component compensates for | Anthropic skills posts; harness ablation practice; skill-creator guidance | Medium |
| 9 | A different-model-family reviewer for high-risk changes such as auth, migrations and irreversible release steps | Self-preference and jury studies, at abstract level | Medium-low; adds cost |
Cost should be budgeted, not feared. Multi-agent systems use about 15 times the tokens of chat (Anthropic, June 2025), but a scoped evaluator was about 8% of Anthropicâs harness spend, so the expensive part of adding judges is the rework they trigger, which is the point. Faber should track tokens and dollars per phase, reserve wide fan-out for phases where checks are cheap and the work is genuinely parallel (research, audits, review, mechanical migrations), and let each validator justify itself through the with/without comparison.
Conclusion
The center of gravity has moved from making agents generate better code to deciding, cheaply and credibly, that the code is done. The components that keep surviving model upgrades (layered criteria, deterministic gates, fresh-context judges with tools, and pass^k evals) all serve that decision, while components that coach the model (context resets, forced decomposition, long system prompts) keep getting deleted as models improve.
That is why the validator is the one piece of a harness unlikely to be erased by better models: as capability grows, teams push work to the new edge, which is exactly where Anthropic found the evaluator âcontinued to give real lift.â For a framework used across many projects, phases are not the differentiator, because every spec-driven tool has them. A portable, measured judge that reads a project-specific contract is.
That reframes Faberâs promise of âearned autonomy,â which its README describes as a progression over days and months. The evidence points to earning autonomy by measurement rather than elapsed time. A project or work type would move from assisted to autonomous when its end-to-end pass^3 and its validatorsâ measured catch rate clear agreed thresholds on its own eval slice, and move back when production rework rises.
The compounding arithmetic shows why calendar-based trust is unsafe: with a generator that succeeds 60% of the time per phase, validators that catch 80% of bad work leave only about half of five-phase runs free of an accepted defect. Faberâs eval suite, run with and without each component, is how Iâll measure what each of these practices is worth on my own work.
This is the eighth post in the Earned Autonomy series and the reference behind the others. Previous: pass^k: The Reliability Math of Unattended Agents. Start at the beginning: Who Decides an Agentâs Work Is Done? How Faber Is Changing.