Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
Abstract
Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). In owner-disjoint folds, one owner half identifies items with low residual differential item functioning across model families (low-DIF). These items are used to score models in the other half with frozen weights that preserve the benchmark's composition across metadata-defined item groups and easiness strata. Equally short matched-random subtests provide a baseline for variation due to item subsampling. Full-benchmark and low-DIF rankings remain strongly correlated (–). Yet in four of five benchmarks, 30.9–47.1% of cross-family pairs initially within one percentage point reverse order, exceeding their matched-random medians by 16.9–28.6 percentage points (all ). The fifth benchmark shows no reliable excess ( points, ). The pattern survives all pre-specified population perturbations, and residual item–family signatures replicate across owner halves; however, no family shows a consistent advantage across benchmarks. Thus, globally stable rankings can still leave individual near-tie orderings sensitive to benchmark composition, and sub-one-point leaderboard gaps should be accompanied by evidence that the implied ordering is composition-robust.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.