acceptodds
Under review as a conference paper at ICLR 2027

HOW STABLE ARE BENCHMARK-DERIVED DIFFICULTY RANKINGS?

Abstract

Analyses of benchmark data often identify which distinctions appear hardest, guiding model diagnosis, data curation, and subsequent evaluation. Yet a difficulty ranking is produced by a measurement procedure: fixed examples are represented, a supervised readout (probe) is fitted, and its scores are ordered. We construct pairwise discrimination tasks from benchmark labels and assess the stability of their ordering under representation and readout choices. On 400 AIOps2025 failure events spanning 55 mechanism pairs, we compare three representations and four fixed probes using grouped out-of-fold evaluation. Across all six probe contrasts under a fixed temporal representation, mean absolute rank shifts range from 3.69 to 14.55 positions; in the motivating logistic-regression–random-forest contrast, only three of the hardest ten pairs overlap. Numerical-feature, fit-seed and alternative grouped-fold controls place this displacement in context. RCAEval RE1 shows heterogeneous probe sensitivity across three distinct systems. Across six OpenML tabular datasets meeting prespecified eligibility criteria, median span-scaled rank displacement ranges from 0.065 to 0.233. Benchmark-derived pair difficulty is a measurable, procedure-relative output whose stability varies across settings and domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.