Earned Autonomy ¡ Part 5
Your Agent Harness Is a List of What the Model Can't Do
Every part of an agent harness is a bet on what the model can't do yet. Which parts go stale at each model upgrade, which never do, and how I'm pruning mine.
Josh McWilliam
10 min read
- agentic-systems
- harness
- production
- claude-code
- platform
The most useful thing a new model does to my harness is break it.
When the newest model generation arrived, three things in Faber turned out to be assumptions Iâd forgotten Iâd made. The Python SDK sent a fixed sampling temperature by default, and the new models reject sampling parameters. Some skills named the model they should run on. Default output limits left no room for the modelâs thinking, which now counts toward them.
None of those were wrong when I wrote them. All three were wrong the day the upgrade landed. Thatâs the idea behind this post: your harness is a list of what the model couldnât do on the day you wrote it.
Every component is a bet
Prithvi Rajasekaran put it in one sentence in Anthropicâs post on harness design for long-running apps (March 2026): âevery component in a harness encodes an assumption about what the model canât do on its own, and those assumptions are worth stress testing, both because they may be incorrect, and because they can quickly go stale as models improve.â
Read your own harness that way and it turns into a list.
- A planner exists because the model under-scopes when itâs handed a raw prompt.
- A context reset exists because the model wraps up early as its context fills.
- A separate evaluator exists because the model praises its own work.
- A retry loop exists because the first attempt fails often enough to matter.
- A long instruction file exists because the model didnât know your conventions.
Every line is a sentence about a weakness. Some of those sentences stop being true. The question is whether you notice when they do.
Watching a harness get thinner
Anthropic has published the clearest record I know of a harness losing weight.
Their long-running coding harness used context resets: clear the window, start a fresh agent, hand over the state in a structured file. The reason, from their Managed Agents post (April 2026): an earlier model âwould wrap up tasks prematurely as it sensed its context limit approachingâa behavior sometimes called âcontext anxiety.ââ
Then they ran the same harness on the next model. âWe found that the behavior was gone. The resets had become dead weight.â
The generation after that took out another piece. The harness had split the build into sprints so the model could stay coherent on one chunk at a time. With a more capable model, Rajasekaran removed the sprint construct entirely and moved the evaluator to a single pass at the end of the run.
He kept the planner, because without it âthe generator under-scoped.â He kept the evaluator too, but found it earned its place only on part of the work. It âis worth the cost when the task sits beyond what the current model does reliably solo.â
The method matters more than the result. His first attempt cut the harness back radically, and âI wasnât able to replicate the performance of the original. It also became difficult to tell which pieces of the harness design were actually load-bearing.â What worked was slower: âremoving one component at a time and reviewing what impact it had on the final result.â
Thatâs an ablation, the oldest move in experimental science, and itâs the right one here. Cut everything and you learn that something mattered. Cut one thing and you learn what.
His conclusion isnât that harnesses disappear. âThe space of interesting harness combinations doesnât shrink as models improve. Instead, it moves.â
How thin can an agent get?
For a sense of the floor, mini-swe-agent is âjust some 100 lines of python for the agent class,â plus a bit more for the environment and the run script. Its only tool is bash. Its maintainers report that it âscores >74% on the SWE-bench verified benchmark.â
Two caveats. A benchmark isnât your codebase. And OpenAI has since explained why SWE-bench Verified no longer measures frontier coding capabilities (February 2026). Its audit found that â59.4% of the 138 problems contained material issues in test design and/or problem description,â and OpenAI stopped reporting the scores.
So read that number as âa very small agent clears most of a benchmark the field has outgrown,â not as a leaderboard. The direction still holds: a lot of what elaborate scaffolds once did, a capable model now does with a shell.
OpenAIâs own harness engineering post (February 2026) shows where the effort went instead. They built and shipped an internal beta of a product with â0 lines of manually-written code.â Their first instinct, one big instruction file, failed. They replaced it with âa short AGENTS.md (roughly 100 lines)â that âserves primarily as a map.â
The instructions got thinner. The environment got thicker: documentation treated as the system of record, custom lints whose error messages âinject remediation instructions into agent context,â and background tasks that scan for drift and open small refactoring pull requests. They call that last part garbage collection.
What goes stale, and what doesnât
Put those examples side by side and two lists fall out.
What goes stale is anything that coaches the model:
- context resets and other workarounds for how a model behaves near its limits
- forced decomposition into small chunks
- long instruction files that restate what a capable model already does
- generic advice in prompts, like âwrite tests alongside codeâ
- pinned sampling settings and pinned model choices
What doesnât is anything that governs the work:
- oracles: tests, compilers, type checks, the build
- security boundaries
- durable state that survives a crash
- budgets and time limits
- isolation between parallel workers
- approval gates for irreversible actions
The first list compensates for what the model canât do. The second would be there even if the model could do everything, because itâs about what you can afford to get wrong.
Anthropicâs Managed Agents post has the cleanest example of the difference. Narrowly scoping a credential the agent can see âencodes an assumption about what Claude canât do with a limited tokenâand Claude is getting increasingly smart.â Their fix was structural: âmake sure the tokens are never reachable from the sandbox where Claudeâs generated code runs.â
A scoping rule bets on the modelâs limits. A boundary doesnât bet at all. As models get more capable, the boundary gets more valuable, not less.
Birgitta BĂśckelerâs article on harness engineering (martinfowler.com, April 2026) gives the two halves of a harness better names. Guides âanticipate the agentâs behaviour and aim to steer it before it acts.â Sensors âobserve after the agent acts and help it self-correct.â
Sheâs also blunt about where sensors stop, whether theyâre deterministic checks or LLM judges: âNeither catches reliably some of the higher-impact problems: Misdiagnosis of issues, overengineering and unnecessary features, misunderstood instructions.â Thatâs why the last thing on the list that never goes stale is a person at the point where the work gets framed. I wrote about where that person sits in the first post in this series.
Skills have two lifespans
Skills split the same way. Anthropicâs skill-creator update (March 2026) sorts them into two kinds. Capability uplift skills âhelp Claude do something the base model either canât do or canât do consistently.â Encoded preference skills âdocument workflows where Claude can already do each piece, but the skill sequences them according to your teamâs process.â
Then the useful part: âCapability uplift skills may become less necessary as models improve. Evals tell you when thatâs happened. Encoded preference skills are more durable.â
It even gives the retirement test: âIf the base model starts passing your evals without the skill loaded, thatâs a signal the skillâs techniques may have been incorporated into the modelâs default behavior. The skill isnât broken; itâs just no longer necessary.â
This site has a good example of the durable kind. Its content guard fails the build if a page contains a claim Iâve decided never to make, a name I donât use, or a link to a page that no longer exists. No model upgrade will make that obsolete, because it isnât compensating for a weakness. Itâs my judgment, written down once. I described it in Vibe Coding Gets You a Demo.
The reliability kit for unattended runs
If coaching goes stale and governance doesnât, the practical question is what governance a run needs when nobody is watching it. Across Anthropicâs harness posts, Nicholas Carliniâs compiler project, OpenAIâs post and the Claude Code docs, the list is short and consistent.
- A fresh session per unit of work, re-grounded from files. In Anthropicâs long-running harness (November 2025), every session starts by reading âthe git logs and progress files to get up to speed on what was recently worked on.â
- Progress in a format the model wonât casually rewrite. The same post: âthe model is less likely to inappropriately change or overwrite JSON files compared to Markdown files.â
- A commit per unit of work, so a bad change can be reverted to a known-good state.
- Caps on turns and dollars, counted across the whole run. The Claude Agent SDKâs
maxBudgetUsdâCounts only the callâs own spend; totals restored from a resumed session donât countâ (TypeScript reference). Its options cap turns and spend, and a wall-clock limit is something you build yourself. A harness that resumes runs needs its own running total and its own clock. - A no-progress stop. The same failure twice, or the same artifact twice, means stop and escalate.
- Deterministic replay. Claude Codeâs workflow runtime makes
Date.now()andMath.random()throw inside scripts âso that a relaunched run repeats the same agent() calls.â - Logs written for an agent to read. From Carliniâs C compiler project (February 2026): âClaude should write ERROR and put the reason on the same line so grep will find it.â He also designed around time blindness. Left alone, the model âwill happily spend hours running tests instead of making progress.â
- Isolation and least privilege. A separate worktree for each parallel writer, and credentials out of reach of generated code.
- Scheduled cleanup, the garbage collection OpenAI describes.
- A result you read, not a status you trust. From Claude Codeâs routines docs: a green status âmeans the session started and exited without an infrastructure error. It does not mean the task in your prompt succeeded.â
Count how many of those ten coach the model. Arguably one, the log format, and even that is about the environment. The rest is infrastructure. Thatâs the part of a harness worth investing in, because itâs the part that survives the next upgrade.
Keep a register
The habit Iâm building from all of this is boring on purpose: a register. One line per harness component, three columns.
- What it compensates for. âThe model under-scopes without a spec.â âThe model praises its own work.â
- How Iâd know itâs no longer needed. Usually the same eval suite, run with and without the component.
- When I last checked. The answer should be âat the last model upgrade.â
Anything I canât fill in the first column for is governance, and it stays. Anything whose second column comes back âpasses without itâ gets removed, one component at a time, the way Rajasekaran did it.
The second column is the hard one. A check that runs inside a pipeline canât tell you whether the pipeline got better. That takes an eval suite, which is the next post.
What this changes in Faber
Faber is the open-source workflow tool the studioâs work runs through, and itâs getting the same treatment. If you want the background on why it exists, itâs in Building the Fractary Platform.
Shipped. When the newest model generation arrived, faber #233 took out three stale assumptions. It removed model names from skills, because skills run inside the main agent and a skill that pins its own model breaks context continuity. It stopped sending a fixed temperature by default, because the new models reject sampling parameters. And it raised default output limits, because thinking now counts toward them. It also re-tiered which model each agent role uses, by how hard the task is.
Planned, in the harness hardening spec:
- Every scaffold records which model weakness it compensates for, and is re-tested at every model upgrade. Thatâs the register, built into the tool.
- The runtime gets a cumulative run budget, because the SDKâs budget doesnât count spend restored on resume, and per-step wall-clock timeouts, so a hung step gets killed instead of waited on.
- An eval pilot on ten past work items, three runs each, supplies the with-and-without comparison the register needs.
The generic coaching in Faberâs prompts goes on the register first. The project standards and the checks stay, because they were never about what the model couldnât do.
This is the fifth post in the Earned Autonomy series. Previous: What Running Thousands of Agents Actually Takes. Next: A Validator Is Not an Eval: How to Measure the Judge.