← Blog

Earned Autonomy ¡ Part 5

Your Agent Harness Is a List of What the Model Can't Do

Every part of an agent harness is a bet on what the model can't do yet. Which parts go stale at each model upgrade, which never do, and how I'm pruning mine.

Josh McWilliam

10 min read

  • agentic-systems
  • harness
  • production
  • claude-code
  • platform

The most useful thing a new model does to my harness is break it.

When the newest model generation arrived, three things in Faber turned out to be assumptions I’d forgotten I’d made. The Python SDK sent a fixed sampling temperature by default, and the new models reject sampling parameters. Some skills named the model they should run on. Default output limits left no room for the model’s thinking, which now counts toward them.

None of those were wrong when I wrote them. All three were wrong the day the upgrade landed. That’s the idea behind this post: your harness is a list of what the model couldn’t do on the day you wrote it.

Every component is a bet

Prithvi Rajasekaran put it in one sentence in Anthropic’s post on harness design for long-running apps (March 2026): “every component in a harness encodes an assumption about what the model can’t do on its own, and those assumptions are worth stress testing, both because they may be incorrect, and because they can quickly go stale as models improve.”

Read your own harness that way and it turns into a list.

  • A planner exists because the model under-scopes when it’s handed a raw prompt.
  • A context reset exists because the model wraps up early as its context fills.
  • A separate evaluator exists because the model praises its own work.
  • A retry loop exists because the first attempt fails often enough to matter.
  • A long instruction file exists because the model didn’t know your conventions.

Every line is a sentence about a weakness. Some of those sentences stop being true. The question is whether you notice when they do.

Watching a harness get thinner

Anthropic has published the clearest record I know of a harness losing weight.

Their long-running coding harness used context resets: clear the window, start a fresh agent, hand over the state in a structured file. The reason, from their Managed Agents post (April 2026): an earlier model “would wrap up tasks prematurely as it sensed its context limit approaching—a behavior sometimes called ‘context anxiety.’”

Then they ran the same harness on the next model. “We found that the behavior was gone. The resets had become dead weight.”

The generation after that took out another piece. The harness had split the build into sprints so the model could stay coherent on one chunk at a time. With a more capable model, Rajasekaran removed the sprint construct entirely and moved the evaluator to a single pass at the end of the run.

He kept the planner, because without it “the generator under-scoped.” He kept the evaluator too, but found it earned its place only on part of the work. It “is worth the cost when the task sits beyond what the current model does reliably solo.”

The method matters more than the result. His first attempt cut the harness back radically, and “I wasn’t able to replicate the performance of the original. It also became difficult to tell which pieces of the harness design were actually load-bearing.” What worked was slower: “removing one component at a time and reviewing what impact it had on the final result.”

That’s an ablation, the oldest move in experimental science, and it’s the right one here. Cut everything and you learn that something mattered. Cut one thing and you learn what.

His conclusion isn’t that harnesses disappear. “The space of interesting harness combinations doesn’t shrink as models improve. Instead, it moves.”

How thin can an agent get?

For a sense of the floor, mini-swe-agent is “just some 100 lines of python for the agent class,” plus a bit more for the environment and the run script. Its only tool is bash. Its maintainers report that it “scores >74% on the SWE-bench verified benchmark.”

Two caveats. A benchmark isn’t your codebase. And OpenAI has since explained why SWE-bench Verified no longer measures frontier coding capabilities (February 2026). Its audit found that “59.4% of the 138 problems contained material issues in test design and/or problem description,” and OpenAI stopped reporting the scores.

So read that number as “a very small agent clears most of a benchmark the field has outgrown,” not as a leaderboard. The direction still holds: a lot of what elaborate scaffolds once did, a capable model now does with a shell.

OpenAI’s own harness engineering post (February 2026) shows where the effort went instead. They built and shipped an internal beta of a product with “0 lines of manually-written code.” Their first instinct, one big instruction file, failed. They replaced it with “a short AGENTS.md (roughly 100 lines)” that “serves primarily as a map.”

The instructions got thinner. The environment got thicker: documentation treated as the system of record, custom lints whose error messages “inject remediation instructions into agent context,” and background tasks that scan for drift and open small refactoring pull requests. They call that last part garbage collection.

What goes stale, and what doesn’t

Put those examples side by side and two lists fall out.

What goes stale is anything that coaches the model:

  • context resets and other workarounds for how a model behaves near its limits
  • forced decomposition into small chunks
  • long instruction files that restate what a capable model already does
  • generic advice in prompts, like “write tests alongside code”
  • pinned sampling settings and pinned model choices

What doesn’t is anything that governs the work:

  • oracles: tests, compilers, type checks, the build
  • security boundaries
  • durable state that survives a crash
  • budgets and time limits
  • isolation between parallel workers
  • approval gates for irreversible actions

The first list compensates for what the model can’t do. The second would be there even if the model could do everything, because it’s about what you can afford to get wrong.

Anthropic’s Managed Agents post has the cleanest example of the difference. Narrowly scoping a credential the agent can see “encodes an assumption about what Claude can’t do with a limited token—and Claude is getting increasingly smart.” Their fix was structural: “make sure the tokens are never reachable from the sandbox where Claude’s generated code runs.”

A scoping rule bets on the model’s limits. A boundary doesn’t bet at all. As models get more capable, the boundary gets more valuable, not less.

Birgitta Böckeler’s article on harness engineering (martinfowler.com, April 2026) gives the two halves of a harness better names. Guides “anticipate the agent’s behaviour and aim to steer it before it acts.” Sensors “observe after the agent acts and help it self-correct.”

She’s also blunt about where sensors stop, whether they’re deterministic checks or LLM judges: “Neither catches reliably some of the higher-impact problems: Misdiagnosis of issues, overengineering and unnecessary features, misunderstood instructions.” That’s why the last thing on the list that never goes stale is a person at the point where the work gets framed. I wrote about where that person sits in the first post in this series.

Skills have two lifespans

Skills split the same way. Anthropic’s skill-creator update (March 2026) sorts them into two kinds. Capability uplift skills “help Claude do something the base model either can’t do or can’t do consistently.” Encoded preference skills “document workflows where Claude can already do each piece, but the skill sequences them according to your team’s process.”

Then the useful part: “Capability uplift skills may become less necessary as models improve. Evals tell you when that’s happened. Encoded preference skills are more durable.”

It even gives the retirement test: “If the base model starts passing your evals without the skill loaded, that’s a signal the skill’s techniques may have been incorporated into the model’s default behavior. The skill isn’t broken; it’s just no longer necessary.”

This site has a good example of the durable kind. Its content guard fails the build if a page contains a claim I’ve decided never to make, a name I don’t use, or a link to a page that no longer exists. No model upgrade will make that obsolete, because it isn’t compensating for a weakness. It’s my judgment, written down once. I described it in Vibe Coding Gets You a Demo.

The reliability kit for unattended runs

If coaching goes stale and governance doesn’t, the practical question is what governance a run needs when nobody is watching it. Across Anthropic’s harness posts, Nicholas Carlini’s compiler project, OpenAI’s post and the Claude Code docs, the list is short and consistent.

  1. A fresh session per unit of work, re-grounded from files. In Anthropic’s long-running harness (November 2025), every session starts by reading “the git logs and progress files to get up to speed on what was recently worked on.”
  2. Progress in a format the model won’t casually rewrite. The same post: “the model is less likely to inappropriately change or overwrite JSON files compared to Markdown files.”
  3. A commit per unit of work, so a bad change can be reverted to a known-good state.
  4. Caps on turns and dollars, counted across the whole run. The Claude Agent SDK’s maxBudgetUsd “Counts only the call’s own spend; totals restored from a resumed session don’t count” (TypeScript reference). Its options cap turns and spend, and a wall-clock limit is something you build yourself. A harness that resumes runs needs its own running total and its own clock.
  5. A no-progress stop. The same failure twice, or the same artifact twice, means stop and escalate.
  6. Deterministic replay. Claude Code’s workflow runtime makes Date.now() and Math.random() throw inside scripts “so that a relaunched run repeats the same agent() calls.”
  7. Logs written for an agent to read. From Carlini’s C compiler project (February 2026): “Claude should write ERROR and put the reason on the same line so grep will find it.” He also designed around time blindness. Left alone, the model “will happily spend hours running tests instead of making progress.”
  8. Isolation and least privilege. A separate worktree for each parallel writer, and credentials out of reach of generated code.
  9. Scheduled cleanup, the garbage collection OpenAI describes.
  10. A result you read, not a status you trust. From Claude Code’s routines docs: a green status “means the session started and exited without an infrastructure error. It does not mean the task in your prompt succeeded.”

Count how many of those ten coach the model. Arguably one, the log format, and even that is about the environment. The rest is infrastructure. That’s the part of a harness worth investing in, because it’s the part that survives the next upgrade.

Keep a register

The habit I’m building from all of this is boring on purpose: a register. One line per harness component, three columns.

  • What it compensates for. “The model under-scopes without a spec.” “The model praises its own work.”
  • How I’d know it’s no longer needed. Usually the same eval suite, run with and without the component.
  • When I last checked. The answer should be “at the last model upgrade.”

Anything I can’t fill in the first column for is governance, and it stays. Anything whose second column comes back “passes without it” gets removed, one component at a time, the way Rajasekaran did it.

The second column is the hard one. A check that runs inside a pipeline can’t tell you whether the pipeline got better. That takes an eval suite, which is the next post.

What this changes in Faber

Faber is the open-source workflow tool the studio’s work runs through, and it’s getting the same treatment. If you want the background on why it exists, it’s in Building the Fractary Platform.

Shipped. When the newest model generation arrived, faber #233 took out three stale assumptions. It removed model names from skills, because skills run inside the main agent and a skill that pins its own model breaks context continuity. It stopped sending a fixed temperature by default, because the new models reject sampling parameters. And it raised default output limits, because thinking now counts toward them. It also re-tiered which model each agent role uses, by how hard the task is.

Planned, in the harness hardening spec:

  • Every scaffold records which model weakness it compensates for, and is re-tested at every model upgrade. That’s the register, built into the tool.
  • The runtime gets a cumulative run budget, because the SDK’s budget doesn’t count spend restored on resume, and per-step wall-clock timeouts, so a hung step gets killed instead of waited on.
  • An eval pilot on ten past work items, three runs each, supplies the with-and-without comparison the register needs.

The generic coaching in Faber’s prompts goes on the register first. The project standards and the checks stay, because they were never about what the model couldn’t do.


This is the fifth post in the Earned Autonomy series. Previous: What Running Thousands of Agents Actually Takes. Next: A Validator Is Not an Eval: How to Measure the Judge.

The Stack

Build on the stack the studio builds with.

Open-source, Apache-2.0 tools for building agentic systems without vendor lock-in. The same stack behind every venture on this site.