Precise Records, Unresolved Ranks: Auditing the Resolution of Agent Leaderboards
Abstract
We audit the SWE-bench Verified leaderboard and find that its ranking means two different things. Read as a record of what happened under the leaderboard's own scoring rule, it is nearly exact: the median compatible-rank width — the number of rank positions still possible for an entry — is 1, and it rises only to 3 when every genuinely unrecoverable outcome is allowed to vary. Read as a measurement of relative performance, it resolves little: simultaneous rank intervals have a median width of 56 positions on a 175-entry leaderboard, and no adjacent pair survives multiple-testing correction. The gap matters to anyone who promotes a system because its score rose. When an acceptance rule keeps the best of several seeds, it promotes all 14 harness variants whose rendered prompts are byte-identical to their baseline — every such promotion is noise. At a budget of 500 tasks, the one-sided fixed-look rule requires a gap of roughly 5.2-6.2 percentage points to reach target power under the discordance envelope measured for one worked pair — several times the gap between adjacent rows. We therefore define the resolution of an evaluation — the smallest effect a declared comparison rule reliably detects at a target power — and release an executable tool that produces certificates binding each reported resolution to its data, rule, assumptions, and provenance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.