acceptodds
Under review as a conference paper at ICLR 2027

Are Agent Reliability Rankings Real? A Statistical Treatment of the Metric

Abstract

The reliability of LLM agents — their tendency to succeed repeatably, not merely once — is increasingly measured by , the probability that an agent succeeds on all independent attempts of a task. It is reported by -bench, -bench and agent-memory benchmarks, and used to rank models. We show that its statistics have never been checked, and that its failures share one mechanism: mis-centering. The ubiquitous plug-in estimator is provably upward biased. We give its exact bias in closed form via Stirling numbers; the relative bias grows without bound in (at , : 3% at , 101% at ). That bias then propagates silently into interval estimates. The hierarchical (task trial) bootstrap resamples within-task counts as , which makes even the unbiased estimator unbiased for rather than : its resampling distribution is centred exactly on the plug-in. Coverage therefore falls as tasks are added — 91% to 36% to 7% at (, ) — which no variance-based account predicts, and re-centring does not repair. We report equally that the plain task-level bootstrap does not fail. We then give the first interval with a finite-sample, distribution-free coverage guarantee, via betting concentration on the bounded unbiased estimator, and compare it against a hierarchical Beta-Binomial baseline. That baseline is narrower where its Beta shape holds; but on a shape agent suites actually have — an atom at , tasks every agent solves — it loses coverage as tasks are added (0.59 at ), the same mis-centering signature, this time from shrinkage rather than resampling. And at the budgets suites run, neither it nor the bootstrap reaches nominal (0.90-0.93 at ), where the guaranteed interval covers 1.00. Inverting the closed-form variance yields a design rule: separation is task-limited, not trial-limited. Finally we test the theory on data. On 4,600 agent rollouts we run ourselves (pre-registered), Theorem 1 predicts the observed bias with no free parameters (correlation 0.969), and the plug-in overstates by 20-147% — worst for the least reliable agent, where deployment is actually in question. A audit of six leaderboards — the regime most favourable to separation, and therefore a lower bound on the problem — finds 33% of all model pairs statistically indistinguishable and 16 of 30 models sharing -bench's top tier. We release a drop-in tool.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.