← Blog

Earned Autonomy Ā· Part 7

pass^k: The Reliability Math of Unattended Agents

Run a 90%-reliable agent ten times and the odds that every run succeeds are about 35%. The arithmetic of trusting agents nobody watches, and what to measure.

Josh McWilliam

9 min read

  • evals
  • harness
  • agentic-systems
  • studio-of-one

A 90% success rate sounds like an agent you can leave alone. Run it ten times, each run independent, and the chance that every run succeeds is about 35%.

That one line changed how I think about unattended work. When nobody is watching, you don’t get to pick the best of several attempts. Every run is the one that ships.

This post is the arithmetic behind that, and what it says about when an agent has earned the right to run without me.

Two numbers that start equal and split apart

Anthropic’s guide to agent evaluation, Demystifying evals for AI agents (January 2026), separates two metrics that are easy to blur:

  • pass@k: the chance that at least one of k attempts succeeds. It rises as k rises. More shots on goal.
  • pass^k: in the guide’s words, ā€œthe probability that all k trials succeed.ā€ It falls as k rises.

The guide’s example: ā€œIf your agent has a 75% per-trial success rate and you run 3 trials, the probability of passing all three is (0.75)³ ā‰ˆ 42%.ā€ And its rule for choosing between them: ā€œpass@k for tools where one success matters, pass^k for agents where consistency is essential.ā€

An agent running overnight is the second kind. A person reviewing five drafts and keeping the best one is buying pass@k. A loop that opens pull requests while I sleep is selling me pass^k, whether or not anyone labels it that way.

Here’s how fast it falls when every run is an independent draw at the same success rate. This is a model, not a measurement:

Per-run successAll 3 runs succeedAll 10 runs succeed
75%42%6%
90%73%35%
95%86%60%
99%97%90%

The bottom row is the uncomfortable one. Even at 99%, an agent that runs every night for ten weeks has roughly even odds of at least one bad night.

Real agents decay more slowly, and that’s the useful part

Real benchmarks don’t fall as fast as that table, because tasks aren’t equally hard.

Sierra’s tau-bench puts an agent in conversation with a simulated user in retail and airline settings, with tools and policies to follow, and reports pass^1 through pass^4. On the retail tasks, the top listed configuration scores 0.692 at pass^1 and 0.462 at pass^4. If every task were equally hard, 0.692 to the fourth power would be about 0.23. The measured number is twice that.

The gap is the shape of the work. Some tasks the agent solves every time. Some it never solves. Some it solves sometimes. As k grows, pass^k converges on the first group: the share of tasks the agent handles every single time.

That changes what I report. A single pass rate hides the split, so I want every task sorted into three buckets:

  • Always. Candidates for running unattended.
  • Never. A capability problem. More harness won’t fix it; a better model or a smaller task might.
  • Sometimes. A reliability problem, and the bucket where my effort goes: sharper acceptance criteria, a check the agent can run, a fresh context on retry.

The ā€œsometimesā€ bucket is where harness work pays. It’s also invisible if you only look at an average.

Phases compound

Most real agent work isn’t one step. A Faber run, for instance, goes through five phases: Frame, Architect, Build, Evaluate, Release, and the work inside them is several steps more.

Reliability multiplies across steps. Five phases that each succeed 95% of the time give you about 77% end to end per run. Ask for three clean runs in a row and you’re near 46%. A model, not a measurement, and a generous one, because it assumes failures are independent.

The lesson isn’t the specific number. It’s where to measure. A dashboard of green per-phase success rates can sit on top of a pipeline that fails one run in four. Measure reliability end to end, not phase by phase.

The judge matters more than the retries

Here’s the part I didn’t expect when I first worked it through.

Take a generator that produces good work 60% of the time. Put a validator after it that passes 90% of good outputs and fails 80% of bad ones. Retry until the validator accepts. Then make the validator stricter at catching bad work, from 80% to 95%, and change nothing else.

Again, a model, not a measurement:

Validator catches 80% of bad workValidator catches 95% of bad work
Accepted outputs that are actually good87%96%
Average attempts per accepted output1.61.8
Five gated phases, no accepted defectabout 50%about 83%

Two different knobs are hiding in a validator, and they do different jobs:

  • How often it catches bad work decides delivered quality. Moving it from 80% to 95% takes a five-phase pipeline from coin-flip to mostly clean.
  • How often it passes good work mostly decides retries and cost. A validator that rejects good work wastes money. A validator that accepts bad work ships defects.

This is why the judge gets its own post in this series, The Agent That Did the Work Shouldn’t Decide It’s Done, and why A Validator Is Not an Eval is about measuring the judge itself. If you don’t know your validator’s catch rate, you don’t know your pipeline’s quality. You know its pass rate, which is a different and more flattering number.

Cost per good result, not cost per run

Cost per run is the number on the invoice. Cost per good result is the number that matters.

Keep the same model and add prices: $2.00 per generation attempt and $0.30 per validation. These are illustrative figures, not anyone’s price list.

  • With the 80% validator, each accepted output costs about $3.71. Since 87% of them are good, each good output costs about $4.26.
  • With the 95% validator, each accepted output costs about $4.11. Since 96% of them are good, each good output costs about $4.26.

The stricter judge costs about forty cents more per run and nothing more per good result. The extra money buys one thing: throwing away bad work you would otherwise have shipped. And that’s before counting the cost of finding a defect later, which in my case means my own hours.

So the cost I track is everything spent (generation, validation, retries, human review) divided by runs that independent checks say were actually correct. Not divided by runs that finished.

This isn’t a new complaint about agent benchmarks. Kapoor and colleagues argued in AI Agents That Matter (July 2024) that with a narrow focus on accuracy, ā€œSOTA agents are needlessly complex and costly,ā€ and proposed ā€œjointly optimizingā€ cost and accuracy instead. The same applies to a harness. A validator that adds a dollar a run and catches nothing is a cost, not a safeguard.

How many tasks before you believe a change

The last piece of arithmetic is the one that keeps me honest about my own improvements.

Anthropic’s evals guide says ā€œ20-50 simple tasks drawn from real failures is a great start,ā€ because ā€œin early agent development, each change to the system often has a clear, noticeable impact, and this large effect size means small sample sizes suffice.ā€ That’s right, and it has a corollary: small suites can only see large effects.

The standard two-proportion calculation (5% significance, 80% power, pass rates near 50%) gives the number of tasks per version you’d need to tell two versions apart. A model, not a measurement:

Real difference between versionsTasks needed per version
30 pointsabout 44
20 pointsabout 98
10 pointsabout 392

With 50 tasks at a pass rate near 50%, the 95% confidence interval is roughly plus or minus 14 points. A validator tweak worth ten points is invisible at that size.

Two things help. First, run both versions on the same tasks. Anthropic’s statistical guidance (November 2024) recommends a paired-differences test because it ā€œlets us eliminate the variance in question difficulty.ā€ If the two versions’ results on each task correlate at around 0.5, that roughly halves the tasks you need. Second, don’t trust naive error bars. The same post notes that ā€œclustered standard errors on popular evals can be over three times as large as naive standard errors.ā€

Running each task more times sharpens the estimate for that task. It doesn’t substitute for more tasks. Thirty runs of ten tasks tell you a lot about ten tasks.

Earned autonomy, by measurement

Faber’s README describes earned autonomy as a calendar. Day 1 is conservative. Week 4 needs less intervention. And then: ā€œMonth 6: Mature—90% autonomous, 10% escalation.ā€

That’s Faber’s own README, and the arithmetic above is why I’m replacing it. Nothing about month six makes a 60% generator with an 80% judge any safer. Under that model, about half of five-phase runs still carry a defect the validator accepted, on day one and on day one hundred and eighty.

The rule I’m moving to is measured, per type of work:

  • A work type, say dependency updates or copy changes, moves from assisted to autonomous when its end-to-end pass^3 and its validators’ measured catch rate clear thresholds I’ve agreed in advance, on its own slice of the eval suite.
  • It moves back when production says otherwise. DORA tracks a metric for exactly that, deployment rework rate: ā€œPercentage of deployments that are unplanned work to fix bugs.ā€
  • New kinds of work start assisted, no matter how long the system has been running.

This matters more when you’re alone. I’ve written about why review is the real cost of agents in a studio of one: I’m the only reviewer there is. Calendar-based trust fails in a specific way for a solo builder. As the system earns trust, volume goes up, my attention per run goes down, and the defects that slip through land on the one person with no one to hand them to.

The point of designing for absence is to be absent safely. That has to be earned with numbers, one work type at a time.

What this changes in Faber

These are planned changes, described in Faber’s harness hardening spec. I’ll mark them shipped here as they land.

  • Autonomy by threshold, not by calendar. Faber’s README already defines autonomy levels, from dry-run through assisted and guarded to autonomous. What changes is how a project moves between them: measured end-to-end pass^3 and validator catch rate, not elapsed time. Planned.
  • An eval pilot before a full suite. Ten past work items, three runs each, on the CLI runtime. The pilot’s numbers set the size of the real suite and how much extra cost a change is allowed to add. Planned.
  • Reporting that matches this post. Pass rates and cost per verified success rather than per run, as the spec sets out, plus my addition: every eval task sorted into always, sometimes and never. Planned.

One piece is already in place: fractary-faber workflow-execute runs the loop as plain code and gives each step a fresh agent session (faber #220). That matters here, because pass^k only means something if the same task can be run the same way three times.

When the pilot has run, I’ll publish the numbers here, including the ones that don’t flatter Faber. The code is on GitHub, and the rest of the stack is on the Platform page.


This is the seventh post in the Earned Autonomy series. Previous: A Validator Is Not an Eval: How to Measure the Judge. Next: Agent Harnesses and Evals: The Full Research Note.

The Stack

Build on the stack the studio builds with.

Open-source, Apache-2.0 tools for building agentic systems without vendor lock-in. The same stack behind every venture on this site.