← Blog

Earned Autonomy ¡ Part 4

What Running Thousands of Agents Actually Takes

The engineer quoted for running thousands of agents still supervises about ten sessions. What scale takes: code that owns the loop and a check that's cheap.

Josh McWilliam

9 min read

  • agentic-systems
  • harness
  • claude-code
  • production

The engineer most often quoted for running thousands of AI agents supervises about ten sessions.

That isn’t a gotcha. It’s the most useful fact in the whole conversation about agent fleets, and almost every retelling drops it. The thousands are real. They just aren’t being watched by a person. They’re being started by loops and fenced by checks that cost almost nothing to run.

If you’re one person trying to get more out of agents, as I am, that distinction is the whole game.

The numbers, read carefully

Boris Cherny created Claude Code, and his numbers climbed all year. Every one of them comes from press coverage of talks and posts, so I’ll attribute each to the outlet that reported it.

In January, VentureBeat reported his setup: “I run 5 Claudes in parallel in my terminal,” plus “5-10 Claudes on claude.ai” in the browser.

In May, Business Insider covered an interview he gave Sequoia Capital. Asked how many sessions he had, he said he typically runs “five to 10 sessions,” each with multiple agents. Then: “Usually, every night, I have like a few thousand that are doing kind of deeper work.” He named the two features he leans on: /loops and Routines, both built for persistent automation.

In June, at Fortune’s Brainstorm Tech, he told the audience: “This morning I was managing maybe a few hundred.” And: “Some days it’s … thousands, or tens of thousands.”

That same month, The New Stack reported that he no longer prompts Claude directly and that “my job is to write loops.” At Meta’s @Scale, TechCrunch quoted him describing a shift “to the point where agents are prompting agents that then write the code,” and calling loops “just as important and as big a step” as the move from source code to agents. His examples, as TechCrunch described them, were standing loops: one agent always looking for ways to improve the code architecture, another hunting for duplicated abstractions to unify.

Now line them up. The number of sessions a human supervises barely moved: ten to fifteen in January, “five to 10” in May. What grew was everything a loop or a schedule started on his behalf.

That’s not a story about attention. It’s a story about automation.

One more line circulates with these quotes: something like “you build the harness that runs the loops.” I went looking for where he said it and couldn’t find a primary source. It reads like a paraphrase of “my job is to write loops,” so I’d quote the original.

Every big fleet stands on a cheap check

Look at the large-scale examples that are documented in detail, and the same thing sits underneath each one: a verifier that is objective, automatic, and nearly free to run.

A C compiler, against GCC. Nicholas Carlini at Anthropic set 16 agents on writing a C compiler in Rust (February 2026). Over nearly 2,000 Claude Code sessions across two weeks, at a total cost just under $20,000, they produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The oracle was GCC, used “as an online known-good compiler oracle to compare against.” His warning is the line I’d frame: “it’s important that the task verifier is nearly perfect.”

A runtime rewrite, against the compiler and the test suite. When Bun moved from Zig to Rust, The Pragmatic Engineer (July 2026) counted 64 agents: four workflow shards, each running 16 Claudes in its own worktree. About 550,000 lines in 11 days. The rewrite finished and nothing compiled, so the compiler’s error output became the work queue, crate by crate, overnight. Then the test suite had to pass in CI. Every commit also went through two adversarial reviews. Sixty-four agents, not thousands, and once the port landed, the work queue came from compiler errors and the test suite rather than from any agent’s opinion.

A browser, against a judge. Cursor ran hundreds of concurrent agents on a single project (Wilson Lin, January 2026). Their first design let agents self-coordinate as equals through a shared file. It failed: “Twenty agents would slow down to the effective throughput of two or three.” Worse, “With no hierarchy, agents became risk-averse.” What worked was planners that create tasks, workers that only execute them, and at the end of each cycle “a judge agent determined whether to continue.”

So scale is a product of two things: orchestration that a script can run, and verification that a machine can do. Take away the second and more agents just means more output waiting for a human to read.

I’ve written about what that review costs a studio of one. It’s the biggest recurring cost I have. Adding agents to a task without a cheap check doesn’t save me hours. It moves them to the end of the week, where they arrive all at once.

”Don’t build agents” never meant fewer agents

Barry Zhang and Mahesh Murag gave a talk at AI Engineer Code 2025 called Don’t Build Agents, Build Skills Instead. It’s easy to hear that as “agents are out.” Read Anthropic’s follow-up write-up (January 2026) and the target is narrower.

“We used to think agents in different domains would look very different,” it says. “A coding agent, a research agent, one for finance, one for marketing—each seemed to need its own tools and scaffolding.” Instead: “Claude Code is a coding agent, but also a general-purpose agent that happens to work through code.” Skills turn “a capable generalist into a knowledgeable specialist.”

What got abandoned was the bespoke agent class per domain, each with hand-built plumbing. What got multiplied was instances of one general harness. The post’s own summary of the stack: “the loop reasons, the runtime executes, MCP connects, and skills guide.”

Four months later the same company launched dynamic workflows (May 2026), in which “Claude dynamically writes orchestration scripts that run tens to hundreds of parallel subagents in a single session, checking its work before anything reaches you.” No contradiction. Fewer kinds of agent, many more runs of the one that works.

Who holds the plan

The clearest way I’ve found to think about this comes from the Claude Code workflow docs. They sort the options by who decides what runs next:

  • Subagents: Claude, turn by turn.
  • Skills: Claude, following the prompt.
  • Agent teams: a lead agent, turn by turn.
  • Workflows: the script.

Then the key sentence: “A workflow moves the plan into code.” With the first three, every result lands in a context window. “A workflow script holds the loop, the branching, and the intermediate results itself, so Claude’s context holds only the final answer.”

There’s a fifth holder the docs cover elsewhere: a schedule. Routines fire on a timer or an event. Even there the docs are careful: “A green status in the run list means the session started and exited without an infrastructure error. It does not mean the task in your prompt succeeded.” A schedule can start the work. It can’t tell you the work is done.

The detail that convinced me this is the right shape is a small one. Inside a workflow script, Claude Code makes Date.now(), Math.random(), and a no-argument new Date() throw, “so that a relaunched run repeats the same agent() calls.” The model writes the orchestration. A deterministic runtime executes it, and can replay it. The runtime also caps the blast radius: up to 16 concurrent agents by default, configurable to 256, and 1,000 agents total per run, which the docs justify in three words: “Prevents runaway loops.”

That’s the twist on the old “workflows versus agents” debate. It isn’t one or the other anymore. The model writes the graph, and code runs it.

To choose between these, the docs on running agents in parallel give three questions: “who coordinates the work, whether the workers need to communicate, and whether they edit the same files.” I now ask them before I start anything with more than one agent in it.

Split the reading, keep one writer

The last question, whether workers edit the same files, is where most multi-agent setups go wrong.

Cognition’s Walden Yan wrote the field’s best-known warning, “Don’t Build Multi-Agents.” In April 2026 he revisited it: the original observations “still hold today for parallel-writer swarms,” but “multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions. A clean-context reviewer catches bugs the coder can’t see.”

The research points the same way. Yubin Kim and colleagues tested 260 configurations across six benchmarks and five architectures. Compared with a single agent, performance ranged “from +80.8% on decomposable financial reasoning to -70.0% on sequential planning,” and “architectures without centralized verification tend to propagate errors more than those with centralized coordination.”

Anthropic’s own multi-agent research system (June 2025) beat a single agent by 90.2% on its internal research eval. It also used “about 15× more tokens than chats,” and the authors flagged that “most coding tasks involve fewer truly parallelizable tasks than research.”

So my rule of thumb is short.

Split when the work is reading (research, search, audits), because a subagent can return what Anthropic calls “a condensed, distilled summary of its work (often 1,000-2,000 tokens)” instead of flooding the main context. Split for independent verification, where a fresh context is the point. Split for per-item transforms that each have their own check and touch their own files, in their own worktree.

Keep one writer when decisions are coupled, when the reasoning is sequential, or when two agents would touch the same file. Parallelism there doesn’t add speed. It adds merge conflicts and quiet disagreements about style that someone, meaning me, has to resolve.

What this means for a studio of one

My attention is the scarce input. It was before agents and it still is. Cherny’s numbers say the same thing about his: around ten sessions he actually watches, no matter how many runs sit underneath.

So the lever was never supervising more agents. It’s writing loops whose output I don’t have to read line by line, because a check already did. On this site that check is one command that runs the type check, the full build, and a content guard that fails the build if any page contains a claim I’ve decided never to make. I described the week it earned its keep in the build log.

The test I now apply before adding agents to anything: is there a cheap, objective way to know each piece is right? If yes, fan out. If no, adding agents just adds reading.

What this changes in Faber

Faber is the workflow tool I build the studio’s work on, and its direction follows directly from this.

Shipped: Faber’s CLI runtime, fractary-faber workflow-execute, already runs the workflow loop as plain code and starts a fresh agent session for each step (faber #220). No master agent decides what runs next.

Merged, in the next release: step verdicts gate the run. A validator’s structured verdict is parsed, and a failing or missing one stops the run (PR #243). In the spec’s words, “Judgment is always a step.”

Planned, in the hardening spec:

  • The CLI runtime becomes the runtime for every run. The chat command becomes a thin wrapper that starts it, relays approvals, and summarizes. The spec’s first principle: “Code runs the loop; agents judge inside steps; humans approve irreversible actions.”
  • Retries and fix loops decided by code. Whether a failed step retries, loops back for a fix, or stops is a rule the runtime applies, not a call an agent makes.
  • Handoffs through files, not memory. In the spec’s words, “Handoffs go through files and version control, never through an agent’s memory.”

Where I want to take it next, from the research note behind this series:

  • Each phase gets a shape. Frame fans out read-only research and has one synthesizer write the acceptance criteria. Architect can draft competing designs, with one decider recording the choice. Build keeps one writer per module, running in parallel only across partitioned units in separate worktrees, with the full test suite as the integration gate. Evaluate runs fresh-context judges, which is the subject of an earlier post in this series.

None of that makes Faber a fleet manager for thousands of agents. It makes it the part that has to exist before thousands would be safe: a loop owned by code, and a judge that isn’t the worker.


This is the fourth post in the Earned Autonomy series. Previous: Done Is a Contract: Four Layers Your Agents Can’t Rewrite. Next: Your Agent Harness Is a List of What the Model Can’t Do.

The Stack

Build on the stack the studio builds with.

Open-source, Apache-2.0 tools for building agentic systems without vendor lock-in. The same stack behind every venture on this site.