acceptodds
Under review as a conference paper at ICLR 2027

What SWE-bench Cannot Tell Apart: The Resolution Limit of Agentic Leaderboards

Abstract

Public leaderboards rank coding agents by resolve rate on a few hundred tasks, and the small gaps at the top drive model selection. We assemble the per-task outcomes of 175 public SWE-bench Verified submissions into a 175x500 matrix that reproduces every official score. The top of the board is compressed: the leading ten systems span 3.4 points, adjacent ranks differ by a median of a single task, and among the top thirty, 243 of the 500 tasks are solved by all of them or by none. Under task subsampling, the 0-1 and 1-2 point gap bands need 497 and 421 of the 500 tasks to reach 95% mean ordering preservation. We then examine sensitivity to run-to-run variation: fifteen pairs of same-agent, same-model resubmissions put the per-run flip rate at 5.6% of tasks (95% CI 3.9 to 7.2). Carrying that rate through the resampling, under an independent equal-probability flip model, leaves the bands within two points below 95% mean ordering accuracy at any subset size, the full benchmark included; at the point estimate 293 of 435 top-thirty pairs (67%) fall short. That result is conditional on the noise model, and a band average does not determine the reliability of any individual pair inside it. A comparison with HELM MMLU shows similar task-sampling behaviour once pool size and score range are matched. Tasks chosen by a held-out panel of comparable systems are worth two to three times their number, but the panel has to be contemporaneous: freezing the chosen set and reusing it on later submissions gives the advantage back. Finally, among the 39 submissions publishing per-task dollars, resolve-rate order and cost-per-resolved-issue order are close to unrelated here (Spearman rho = -0.08).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.