acceptodds
Under review as a conference paper at ICLR 2027

WHEN DOES A RANDOM SUBSET PRESERVE THE FULL-BENCHMARK MODEL RANKING?

Abstract

Evaluating large language models (LLMs) is expensive, so a natural strategy is to evaluate each model on only a random subset of benchmark questions. But how many questions are enough? We show that subset size alone is a poor guide: at the same evaluation budget, the agreement between the subset ranking and the full-benchmark ranking can vary substantially across benchmarks. The missing factor is how informative benchmark questions are for distinguishing the particular models being compared. We characterize this using population-relative item discrimination and combine it with classical measurement theory to predict how closely a random subset recovers the full-benchmark model ranking. Across 114 benchmark instances from 11 suites and five evaluation budgets spanning 1%–20%, predicted and observed ranking agreement correlate at , with correlations ranging from to when evaluated separately at each budget. The predictor also estimates the subset size required for a target level of ranking agreement and the potential for further improvements from evaluating more questions. Finally, we show that achievable ranking agreement depends on the model population itself: when evaluated models become more similar, benchmarks can become substantially less informative, and discrimination estimates from a mismatched source population may not transfer. Overall, our results provide a simple way to assess when random sampling can recover full-benchmark rankings, how many questions are needed, and when benchmark properties limit how well a subset can recover the full ranking.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.