acceptodds
Under review as a conference paper at ICLR 2027

Population-Conditioned Benchmark Item Influence Beyond Difficulty and Discrimination

Abstract

Benchmark item importance is often summarized by pass rate and discrimination. For model ranking, however, these summaries leave a population-conditioned component: which items matter depends on which models are being compared. We introduce population-conditioned ranking influence, a close-pair leverage statistic measuring which items distinguish similarly performing models within a specified population. Across nine benchmark matrices, residual fingerprints retain reproducible population-specific structure after controlling for population-local pass rate, pass rate squared, and corrected item–total discrimination, including under target-excluded capability definitions. On an exactly aligned 441-model by 12,032-item MMLU-Pro matrix, fully disjoint panels independently recover this structure: increasing panel size from 75 to 100 raises median within-population Spearman similarity from 0.361 to 0.435, while cross-population similarity remains much smaller, changing from 0.072 to 0.081. Fixed-margin Curveball randomization does not reproduce this agreement. The same disjoint-panel ordering appears across four additional large response matrices spanning scientific questions, multi-step reasoning, instruction following, and hard mathematics, with GPQA providing a second strong fixed-margin-null-confirmed result. Finally, signed item sets learned from one capability population preserve close-pair rankings of unseen models from that population better than sets learned from mismatched populations; all six paired contrasts are positive with 95% confidence intervals excluding zero. Together, these results show that benchmark ranking importance is relational rather than intrinsic: which questions matter depends on which models are being distinguished, and that population-specific structure predicts relative item-set utility on unseen models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.