HOW STABLE ARE BENCHMARK-DERIVED DIFFICULTY RANKINGS?
Abstract
Analyses of benchmark data often identify which distinctions appear hardest, guiding model diagnosis, data curation, and subsequent evaluation. Yet a difficulty ranking is produced by a measurement procedure: fixed examples are represented, a supervised readout (probe) is fitted, and its scores are ordered. We construct pairwise discrimination tasks from benchmark labels and assess the stability of their ordering under representation and readout choices. On 400 AIOps2025 failure events spanning 55 mechanism pairs, we compare three representations and four fixed probes using grouped out-of-fold evaluation. Across all six probe contrasts under a fixed temporal representation, mean absolute rank shifts range from 3.69 to 14.55 positions; in the motivating logistic-regression–random-forest contrast, only three of the hardest ten pairs overlap. Numerical-feature, fit-seed and alternative grouped-fold controls place this displacement in context. RCAEval RE1 shows heterogeneous probe sensitivity across three distinct systems. Across six OpenML tabular datasets meeting prespecified eligibility criteria, median span-scaled rank displacement ranges from 0.065 to 0.233. Benchmark-derived pair difficulty is a measurable, procedure-relative output whose stability varies across settings and domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.