Earned Autonomy Ā· Part 7
pass^k: The Reliability Math of Unattended Agents
Run a 90%-reliable agent ten times and the odds that every run succeeds are about 35%. The arithmetic of trusting agents nobody watches, and what to measure.
Josh McWilliam
9 min read
- evals
- harness
- agentic-systems
- studio-of-one
A 90% success rate sounds like an agent you can leave alone. Run it ten times, each run independent, and the chance that every run succeeds is about 35%.
That one line changed how I think about unattended work. When nobody is watching, you donāt get to pick the best of several attempts. Every run is the one that ships.
This post is the arithmetic behind that, and what it says about when an agent has earned the right to run without me.
Two numbers that start equal and split apart
Anthropicās guide to agent evaluation, Demystifying evals for AI agents (January 2026), separates two metrics that are easy to blur:
- pass@k: the chance that at least one of k attempts succeeds. It rises as k rises. More shots on goal.
- pass^k: in the guideās words, āthe probability that all k trials succeed.ā It falls as k rises.
The guideās example: āIf your agent has a 75% per-trial success rate and you run 3 trials, the probability of passing all three is (0.75)³ ā 42%.ā And its rule for choosing between them: āpass@k for tools where one success matters, pass^k for agents where consistency is essential.ā
An agent running overnight is the second kind. A person reviewing five drafts and keeping the best one is buying pass@k. A loop that opens pull requests while I sleep is selling me pass^k, whether or not anyone labels it that way.
Hereās how fast it falls when every run is an independent draw at the same success rate. This is a model, not a measurement:
| Per-run success | All 3 runs succeed | All 10 runs succeed |
|---|---|---|
| 75% | 42% | 6% |
| 90% | 73% | 35% |
| 95% | 86% | 60% |
| 99% | 97% | 90% |
The bottom row is the uncomfortable one. Even at 99%, an agent that runs every night for ten weeks has roughly even odds of at least one bad night.
Real agents decay more slowly, and thatās the useful part
Real benchmarks donāt fall as fast as that table, because tasks arenāt equally hard.
Sierraās tau-bench puts an agent in conversation with a simulated user in retail and airline settings, with tools and policies to follow, and reports pass^1 through pass^4. On the retail tasks, the top listed configuration scores 0.692 at pass^1 and 0.462 at pass^4. If every task were equally hard, 0.692 to the fourth power would be about 0.23. The measured number is twice that.
The gap is the shape of the work. Some tasks the agent solves every time. Some it never solves. Some it solves sometimes. As k grows, pass^k converges on the first group: the share of tasks the agent handles every single time.
That changes what I report. A single pass rate hides the split, so I want every task sorted into three buckets:
- Always. Candidates for running unattended.
- Never. A capability problem. More harness wonāt fix it; a better model or a smaller task might.
- Sometimes. A reliability problem, and the bucket where my effort goes: sharper acceptance criteria, a check the agent can run, a fresh context on retry.
The āsometimesā bucket is where harness work pays. Itās also invisible if you only look at an average.
Phases compound
Most real agent work isnāt one step. A Faber run, for instance, goes through five phases: Frame, Architect, Build, Evaluate, Release, and the work inside them is several steps more.
Reliability multiplies across steps. Five phases that each succeed 95% of the time give you about 77% end to end per run. Ask for three clean runs in a row and youāre near 46%. A model, not a measurement, and a generous one, because it assumes failures are independent.
The lesson isnāt the specific number. Itās where to measure. A dashboard of green per-phase success rates can sit on top of a pipeline that fails one run in four. Measure reliability end to end, not phase by phase.
The judge matters more than the retries
Hereās the part I didnāt expect when I first worked it through.
Take a generator that produces good work 60% of the time. Put a validator after it that passes 90% of good outputs and fails 80% of bad ones. Retry until the validator accepts. Then make the validator stricter at catching bad work, from 80% to 95%, and change nothing else.
Again, a model, not a measurement:
| Validator catches 80% of bad work | Validator catches 95% of bad work | |
|---|---|---|
| Accepted outputs that are actually good | 87% | 96% |
| Average attempts per accepted output | 1.6 | 1.8 |
| Five gated phases, no accepted defect | about 50% | about 83% |
Two different knobs are hiding in a validator, and they do different jobs:
- How often it catches bad work decides delivered quality. Moving it from 80% to 95% takes a five-phase pipeline from coin-flip to mostly clean.
- How often it passes good work mostly decides retries and cost. A validator that rejects good work wastes money. A validator that accepts bad work ships defects.
This is why the judge gets its own post in this series, The Agent That Did the Work Shouldnāt Decide Itās Done, and why A Validator Is Not an Eval is about measuring the judge itself. If you donāt know your validatorās catch rate, you donāt know your pipelineās quality. You know its pass rate, which is a different and more flattering number.
Cost per good result, not cost per run
Cost per run is the number on the invoice. Cost per good result is the number that matters.
Keep the same model and add prices: $2.00 per generation attempt and $0.30 per validation. These are illustrative figures, not anyoneās price list.
- With the 80% validator, each accepted output costs about $3.71. Since 87% of them are good, each good output costs about $4.26.
- With the 95% validator, each accepted output costs about $4.11. Since 96% of them are good, each good output costs about $4.26.
The stricter judge costs about forty cents more per run and nothing more per good result. The extra money buys one thing: throwing away bad work you would otherwise have shipped. And thatās before counting the cost of finding a defect later, which in my case means my own hours.
So the cost I track is everything spent (generation, validation, retries, human review) divided by runs that independent checks say were actually correct. Not divided by runs that finished.
This isnāt a new complaint about agent benchmarks. Kapoor and colleagues argued in AI Agents That Matter (July 2024) that with a narrow focus on accuracy, āSOTA agents are needlessly complex and costly,ā and proposed ājointly optimizingā cost and accuracy instead. The same applies to a harness. A validator that adds a dollar a run and catches nothing is a cost, not a safeguard.
How many tasks before you believe a change
The last piece of arithmetic is the one that keeps me honest about my own improvements.
Anthropicās evals guide says ā20-50 simple tasks drawn from real failures is a great start,ā because āin early agent development, each change to the system often has a clear, noticeable impact, and this large effect size means small sample sizes suffice.ā Thatās right, and it has a corollary: small suites can only see large effects.
The standard two-proportion calculation (5% significance, 80% power, pass rates near 50%) gives the number of tasks per version youād need to tell two versions apart. A model, not a measurement:
| Real difference between versions | Tasks needed per version |
|---|---|
| 30 points | about 44 |
| 20 points | about 98 |
| 10 points | about 392 |
With 50 tasks at a pass rate near 50%, the 95% confidence interval is roughly plus or minus 14 points. A validator tweak worth ten points is invisible at that size.
Two things help. First, run both versions on the same tasks. Anthropicās statistical guidance (November 2024) recommends a paired-differences test because it ālets us eliminate the variance in question difficulty.ā If the two versionsā results on each task correlate at around 0.5, that roughly halves the tasks you need. Second, donāt trust naive error bars. The same post notes that āclustered standard errors on popular evals can be over three times as large as naive standard errors.ā
Running each task more times sharpens the estimate for that task. It doesnāt substitute for more tasks. Thirty runs of ten tasks tell you a lot about ten tasks.
Earned autonomy, by measurement
Faberās README describes earned autonomy as a calendar. Day 1 is conservative. Week 4 needs less intervention. And then: āMonth 6: Matureā90% autonomous, 10% escalation.ā
Thatās Faberās own README, and the arithmetic above is why Iām replacing it. Nothing about month six makes a 60% generator with an 80% judge any safer. Under that model, about half of five-phase runs still carry a defect the validator accepted, on day one and on day one hundred and eighty.
The rule Iām moving to is measured, per type of work:
- A work type, say dependency updates or copy changes, moves from assisted to autonomous when its end-to-end pass^3 and its validatorsā measured catch rate clear thresholds Iāve agreed in advance, on its own slice of the eval suite.
- It moves back when production says otherwise. DORA tracks a metric for exactly that, deployment rework rate: āPercentage of deployments that are unplanned work to fix bugs.ā
- New kinds of work start assisted, no matter how long the system has been running.
This matters more when youāre alone. Iāve written about why review is the real cost of agents in a studio of one: Iām the only reviewer there is. Calendar-based trust fails in a specific way for a solo builder. As the system earns trust, volume goes up, my attention per run goes down, and the defects that slip through land on the one person with no one to hand them to.
The point of designing for absence is to be absent safely. That has to be earned with numbers, one work type at a time.
What this changes in Faber
These are planned changes, described in Faberās harness hardening spec. Iāll mark them shipped here as they land.
- Autonomy by threshold, not by calendar. Faberās README already defines autonomy levels, from
dry-runthroughassistedandguardedtoautonomous. What changes is how a project moves between them: measured end-to-end pass^3 and validator catch rate, not elapsed time. Planned. - An eval pilot before a full suite. Ten past work items, three runs each, on the CLI runtime. The pilotās numbers set the size of the real suite and how much extra cost a change is allowed to add. Planned.
- Reporting that matches this post. Pass rates and cost per verified success rather than per run, as the spec sets out, plus my addition: every eval task sorted into always, sometimes and never. Planned.
One piece is already in place: fractary-faber workflow-execute runs the loop as plain code and gives each step a fresh agent session (faber #220). That matters here, because pass^k only means something if the same task can be run the same way three times.
When the pilot has run, Iāll publish the numbers here, including the ones that donāt flatter Faber. The code is on GitHub, and the rest of the stack is on the Platform page.
This is the seventh post in the Earned Autonomy series. Previous: A Validator Is Not an Eval: How to Measure the Judge. Next: Agent Harnesses and Evals: The Full Research Note.