acceptodds
Under review as a conference paper at ICLR 2027

Precise Records, Unresolved Ranks: Auditing the Resolution of Agent Leaderboards

Abstract

We audit the SWE-bench Verified leaderboard and find that its ranking means two different things. Read as a record of what happened under the leaderboard's own scoring rule, it is nearly exact: the median compatible-rank width — the number of rank positions still possible for an entry — is 1, and it rises only to 3 when every genuinely unrecoverable outcome is allowed to vary. Read as a measurement of relative performance, it resolves little: simultaneous rank intervals have a median width of 56 positions on a 175-entry leaderboard, and no adjacent pair survives multiple-testing correction. The gap matters to anyone who promotes a system because its score rose. When an acceptance rule keeps the best of several seeds, it promotes all 14 harness variants whose rendered prompts are byte-identical to their baseline — every such promotion is noise. At a budget of 500 tasks, the one-sided fixed-look rule requires a gap of roughly 5.2-6.2 percentage points to reach target power under the discordance envelope measured for one worked pair — several times the gap between adjacent rows. We therefore define the resolution of an evaluation — the smallest effect a declared comparison rule reliably detects at a target power — and release an executable tool that produces certificates binding each reported resolution to its data, rule, assumptions, and provenance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.