acceptodds
Under review as a conference paper at ICLR 2027

When Are Smaller Benchmarks Better?

Abstract

Automated evaluations often rank systems differently from humans. We show that agreement improves substantially when the benchmark is reduced to a small subset of its existing items, without adding new data. Sieve selects as small as 20 items on a Pareto front over two objectives: per-item accuracy match with the source benchmark, and correlation with the human system ranking. We evaluate on held-out systems and held-out items, so the gain cannot be circular. The 20-item subsets agree with humans better than the full 80- to 5,000-item benchmarks they came from, closing about half the remaining gap to perfect agreement on average and all of it for the weakest judges. Across the 36-cell task-judge grid the mean gain is +0.37 at one-sided Wilcoxon p=2.7×10⁻⁶. At the extreme, LLaMA-3B on MT-Bench swings from an inverted ranking, ρ_T=−0.49, to perfect agreement, ρ_Q*=+1.00. What a judge gains, it turns out, depends on how much room it had to improve. We call this the Headroom Law: gain scales with headroom at slope α=0.86, 95% CI [0.75, 0.93], with per-judge slopes in [0.65, 0.94] and seven independent selectors recovering α ∈ [0.87, 1.01]. So we can tell whether Sieve will help on a new task-judge pair before running the search at all: leave-one-out r=0.92 on fit cells, and live prediction r=0.935 on 66 cells we had never seen.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.