When Are Smaller Benchmarks Better?
Abstract
Automated evaluations often rank systems differently from humans. We show that agreement improves substantially when the benchmark is reduced to a small subset of its existing items, without adding new data. Sieve selects as small as 20 items on a Pareto front over two objectives: per-item accuracy match with the source benchmark, and correlation with the human system ranking. We evaluate on held-out systems and held-out items, so the gain cannot be circular. The 20-item subsets agree with humans better than the full 80- to 5,000-item benchmarks they came from, closing about half the remaining gap to perfect agreement on average and all of it for the weakest judges. Across the 36-cell task-judge grid the mean gain is +0.37 at one-sided Wilcoxon p=2.7×10⁻⁶. At the extreme, LLaMA-3B on MT-Bench swings from an inverted ranking, ρ_T=−0.49, to perfect agreement, ρ_Q*=+1.00. What a judge gains, it turns out, depends on how much room it had to improve. We call this the Headroom Law: gain scales with headroom at slope α=0.86, 95% CI [0.75, 0.93], with per-judge slopes in [0.65, 0.94] and seven independent selectors recovering α ∈ [0.87, 1.01]. So we can tell whether Sieve will help on a new task-judge pair before running the search at all: leave-one-out r=0.92 on fit cells, and live prediction r=0.935 on 66 cells we had never seen.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.