Can Saturated Benchmarks Be Recovered by LLM-as-a-Meta-Judge?
Abstract
Widely used language-model benchmarks are increasingly saturated, with frontier systems receiving near-tied scores that standard metrics cannot resolve. Rather than constructing harder alternatives, we ask whether existing benchmarks can regain informativeness through improved evaluation of the same candidate outputs. Therefore, we present Seeded Elimination with Adaptive LLM-as-a-Meta-Judge (SEAL), which combines a seeded tournament with evolving checklist criteria anchored to fixed task-level principles. To validate the recovered rankings, we construct a human-verified dense reference (HVR) by auditing exhaustive pairwise judgment traces against task and response evidence. Experiments span multiple saturated benchmarks such as code generation, mathematical reasoning, knowledge-intensive question answering, and tool calling. SEAL achieves the highest or joint-highest Spearman agreement with HVR across all settings, averaging 0.733 compared with 0.623 for exhaustive pairwise judging. On code generation, it uses 57.5% fewer judge calls and 65.8% fewer tokens than exhaustive comparison. Beyond the tournament, reusing evolved rubrics in pointwise scoring improves HVR agreement, raising mean correlation from 0.565 to 0.647. Together, these results demonstrate its effectiveness in adapting both comparisons and criteria to recover task-grounded ranking signal from saturated benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.