Earned Autonomy ¡ Part 4
What Running Thousands of Agents Actually Takes
The engineer quoted for running thousands of agents still supervises about ten sessions. What scale takes: code that owns the loop and a check that's cheap.
Josh McWilliam
9 min read
- agentic-systems
- harness
- claude-code
- production
The engineer most often quoted for running thousands of AI agents supervises about ten sessions.
That isnât a gotcha. Itâs the most useful fact in the whole conversation about agent fleets, and almost every retelling drops it. The thousands are real. They just arenât being watched by a person. Theyâre being started by loops and fenced by checks that cost almost nothing to run.
If youâre one person trying to get more out of agents, as I am, that distinction is the whole game.
The numbers, read carefully
Boris Cherny created Claude Code, and his numbers climbed all year. Every one of them comes from press coverage of talks and posts, so Iâll attribute each to the outlet that reported it.
In January, VentureBeat reported his setup: âI run 5 Claudes in parallel in my terminal,â plus â5-10 Claudes on claude.aiâ in the browser.
In May, Business Insider covered an interview he gave Sequoia Capital. Asked how many sessions he had, he said he typically runs âfive to 10 sessions,â each with multiple agents. Then: âUsually, every night, I have like a few thousand that are doing kind of deeper work.â He named the two features he leans on: /loops and Routines, both built for persistent automation.
In June, at Fortuneâs Brainstorm Tech, he told the audience: âThis morning I was managing maybe a few hundred.â And: âSome days itâs ⌠thousands, or tens of thousands.â
That same month, The New Stack reported that he no longer prompts Claude directly and that âmy job is to write loops.â At Metaâs @Scale, TechCrunch quoted him describing a shift âto the point where agents are prompting agents that then write the code,â and calling loops âjust as important and as big a stepâ as the move from source code to agents. His examples, as TechCrunch described them, were standing loops: one agent always looking for ways to improve the code architecture, another hunting for duplicated abstractions to unify.
Now line them up. The number of sessions a human supervises barely moved: ten to fifteen in January, âfive to 10â in May. What grew was everything a loop or a schedule started on his behalf.
Thatâs not a story about attention. Itâs a story about automation.
One more line circulates with these quotes: something like âyou build the harness that runs the loops.â I went looking for where he said it and couldnât find a primary source. It reads like a paraphrase of âmy job is to write loops,â so Iâd quote the original.
Every big fleet stands on a cheap check
Look at the large-scale examples that are documented in detail, and the same thing sits underneath each one: a verifier that is objective, automatic, and nearly free to run.
A C compiler, against GCC. Nicholas Carlini at Anthropic set 16 agents on writing a C compiler in Rust (February 2026). Over nearly 2,000 Claude Code sessions across two weeks, at a total cost just under $20,000, they produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The oracle was GCC, used âas an online known-good compiler oracle to compare against.â His warning is the line Iâd frame: âitâs important that the task verifier is nearly perfect.â
A runtime rewrite, against the compiler and the test suite. When Bun moved from Zig to Rust, The Pragmatic Engineer (July 2026) counted 64 agents: four workflow shards, each running 16 Claudes in its own worktree. About 550,000 lines in 11 days. The rewrite finished and nothing compiled, so the compilerâs error output became the work queue, crate by crate, overnight. Then the test suite had to pass in CI. Every commit also went through two adversarial reviews. Sixty-four agents, not thousands, and once the port landed, the work queue came from compiler errors and the test suite rather than from any agentâs opinion.
A browser, against a judge. Cursor ran hundreds of concurrent agents on a single project (Wilson Lin, January 2026). Their first design let agents self-coordinate as equals through a shared file. It failed: âTwenty agents would slow down to the effective throughput of two or three.â Worse, âWith no hierarchy, agents became risk-averse.â What worked was planners that create tasks, workers that only execute them, and at the end of each cycle âa judge agent determined whether to continue.â
So scale is a product of two things: orchestration that a script can run, and verification that a machine can do. Take away the second and more agents just means more output waiting for a human to read.
Iâve written about what that review costs a studio of one. Itâs the biggest recurring cost I have. Adding agents to a task without a cheap check doesnât save me hours. It moves them to the end of the week, where they arrive all at once.
âDonât build agentsâ never meant fewer agents
Barry Zhang and Mahesh Murag gave a talk at AI Engineer Code 2025 called Donât Build Agents, Build Skills Instead. Itâs easy to hear that as âagents are out.â Read Anthropicâs follow-up write-up (January 2026) and the target is narrower.
âWe used to think agents in different domains would look very different,â it says. âA coding agent, a research agent, one for finance, one for marketingâeach seemed to need its own tools and scaffolding.â Instead: âClaude Code is a coding agent, but also a general-purpose agent that happens to work through code.â Skills turn âa capable generalist into a knowledgeable specialist.â
What got abandoned was the bespoke agent class per domain, each with hand-built plumbing. What got multiplied was instances of one general harness. The postâs own summary of the stack: âthe loop reasons, the runtime executes, MCP connects, and skills guide.â
Four months later the same company launched dynamic workflows (May 2026), in which âClaude dynamically writes orchestration scripts that run tens to hundreds of parallel subagents in a single session, checking its work before anything reaches you.â No contradiction. Fewer kinds of agent, many more runs of the one that works.
Who holds the plan
The clearest way Iâve found to think about this comes from the Claude Code workflow docs. They sort the options by who decides what runs next:
- Subagents: Claude, turn by turn.
- Skills: Claude, following the prompt.
- Agent teams: a lead agent, turn by turn.
- Workflows: the script.
Then the key sentence: âA workflow moves the plan into code.â With the first three, every result lands in a context window. âA workflow script holds the loop, the branching, and the intermediate results itself, so Claudeâs context holds only the final answer.â
Thereâs a fifth holder the docs cover elsewhere: a schedule. Routines fire on a timer or an event. Even there the docs are careful: âA green status in the run list means the session started and exited without an infrastructure error. It does not mean the task in your prompt succeeded.â A schedule can start the work. It canât tell you the work is done.
The detail that convinced me this is the right shape is a small one. Inside a workflow script, Claude Code makes Date.now(), Math.random(), and a no-argument new Date() throw, âso that a relaunched run repeats the same agent() calls.â The model writes the orchestration. A deterministic runtime executes it, and can replay it. The runtime also caps the blast radius: up to 16 concurrent agents by default, configurable to 256, and 1,000 agents total per run, which the docs justify in three words: âPrevents runaway loops.â
Thatâs the twist on the old âworkflows versus agentsâ debate. It isnât one or the other anymore. The model writes the graph, and code runs it.
To choose between these, the docs on running agents in parallel give three questions: âwho coordinates the work, whether the workers need to communicate, and whether they edit the same files.â I now ask them before I start anything with more than one agent in it.
Split the reading, keep one writer
The last question, whether workers edit the same files, is where most multi-agent setups go wrong.
Cognitionâs Walden Yan wrote the fieldâs best-known warning, âDonât Build Multi-Agents.â In April 2026 he revisited it: the original observations âstill hold today for parallel-writer swarms,â but âmulti-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions. A clean-context reviewer catches bugs the coder canât see.â
The research points the same way. Yubin Kim and colleagues tested 260 configurations across six benchmarks and five architectures. Compared with a single agent, performance ranged âfrom +80.8% on decomposable financial reasoning to -70.0% on sequential planning,â and âarchitectures without centralized verification tend to propagate errors more than those with centralized coordination.â
Anthropicâs own multi-agent research system (June 2025) beat a single agent by 90.2% on its internal research eval. It also used âabout 15Ă more tokens than chats,â and the authors flagged that âmost coding tasks involve fewer truly parallelizable tasks than research.â
So my rule of thumb is short.
Split when the work is reading (research, search, audits), because a subagent can return what Anthropic calls âa condensed, distilled summary of its work (often 1,000-2,000 tokens)â instead of flooding the main context. Split for independent verification, where a fresh context is the point. Split for per-item transforms that each have their own check and touch their own files, in their own worktree.
Keep one writer when decisions are coupled, when the reasoning is sequential, or when two agents would touch the same file. Parallelism there doesnât add speed. It adds merge conflicts and quiet disagreements about style that someone, meaning me, has to resolve.
What this means for a studio of one
My attention is the scarce input. It was before agents and it still is. Chernyâs numbers say the same thing about his: around ten sessions he actually watches, no matter how many runs sit underneath.
So the lever was never supervising more agents. Itâs writing loops whose output I donât have to read line by line, because a check already did. On this site that check is one command that runs the type check, the full build, and a content guard that fails the build if any page contains a claim Iâve decided never to make. I described the week it earned its keep in the build log.
The test I now apply before adding agents to anything: is there a cheap, objective way to know each piece is right? If yes, fan out. If no, adding agents just adds reading.
What this changes in Faber
Faber is the workflow tool I build the studioâs work on, and its direction follows directly from this.
Shipped: Faberâs CLI runtime, fractary-faber workflow-execute, already runs the workflow loop as plain code and starts a fresh agent session for each step (faber #220). No master agent decides what runs next.
Merged, in the next release: step verdicts gate the run. A validatorâs structured verdict is parsed, and a failing or missing one stops the run (PR #243). In the specâs words, âJudgment is always a step.â
Planned, in the hardening spec:
- The CLI runtime becomes the runtime for every run. The chat command becomes a thin wrapper that starts it, relays approvals, and summarizes. The specâs first principle: âCode runs the loop; agents judge inside steps; humans approve irreversible actions.â
- Retries and fix loops decided by code. Whether a failed step retries, loops back for a fix, or stops is a rule the runtime applies, not a call an agent makes.
- Handoffs through files, not memory. In the specâs words, âHandoffs go through files and version control, never through an agentâs memory.â
Where I want to take it next, from the research note behind this series:
- Each phase gets a shape. Frame fans out read-only research and has one synthesizer write the acceptance criteria. Architect can draft competing designs, with one decider recording the choice. Build keeps one writer per module, running in parallel only across partitioned units in separate worktrees, with the full test suite as the integration gate. Evaluate runs fresh-context judges, which is the subject of an earlier post in this series.
None of that makes Faber a fleet manager for thousands of agents. It makes it the part that has to exist before thousands would be safe: a loop owned by code, and a judge that isnât the worker.
This is the fourth post in the Earned Autonomy series. Previous: Done Is a Contract: Four Layers Your Agents Canât Rewrite. Next: Your Agent Harness Is a List of What the Model Canât Do.