acceptodds
Under review as a conference paper at ICLR 2027

When Do Harder Tasks Preserve Agent Rankings?

Abstract

Adding harder tasks is a natural way to refresh an agent benchmark, but it can reduce rather than improve resolution when strong systems fail those tasks together. We introduce measurement-regime diagnosis, a post-pilot audit that tests whether hardness agrees with finite-budget ranking recovery and frontier-local distinguishability. On SWE-bench Verified, the 50 hardest and frontier-informative tasks are disjoint, and 96% of the hard set is all-fail among frontier systems. On Terminal-Bench 2, the same diagnostics largely retain hardness as a candidate, although the alignment is outcome-encoding sensitive. SWE-bench transfer changes ordering with budget. Highest frontier information (HFI) has higher fixed-reference correlation than hardness at 25 and 50 tasks, while hardness catches up at 100. None improves the anchor. The pattern survives 10 task splits, repository partitions, and budgets matched on pre-cutoff estimated API calls. At , ranges cross zero, and one family omission changes the ordering. Across four Terminal-Bench 2 chronology/missingness variants, HFI-minus-hardest recovery is positive, but the best policy changes. We therefore formulate a regime-, target-, and budget-indexed Pareto decision with abstention. The evidence is retrospective and conditional on observed systems and piloted candidate outcomes. It does not establish zero-shot utility for newly authored tasks or an absolute benchmark ceiling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.