The Wrong Kind of Benchmark Difficulty: Diagnosing Benchmark Errors Using Model Disagreement
Abstract
Current AI benchmarking practice has aimed for harder and harder questions. Yet increasingly, evidence suggests that hard questions may just be fallacious ones. We offer and test the reliability of a remarkably simple diagnostic approach for surfacing such questions: expanding answer choices to allow abstention or multiple correct answers and assessing model disagreement (from keyed answer and convergence on the same non-keyed answer) for frontier models. First, we show that a panel as small as , with full disagreement and any amount of convergence, recovers the per-subject quality profile established by previously published, independent expert audits of two popular benchmarks, Humanity's Last Exam (HLE) and Massive Multitask Language Understanding (MMLU), without per-benchmark tuning. Second, we demonstrate that this simple three-model disagreement signal outperforms prior disagreement-based screening and an LLM-as-judge panel, recovering independently verified errors with 72–84% precision and 61–65% recall. Third, we assess key assumptions, including independence, and provide evidence that model diversity enhances performance more so than panel size. Fourth, we highlight the recency of these diagnostic capabilities by running the same procedure with earlier model generations and tracing improvements, particularly in precision, across successive model generations. Fifth, we show that the current trend to design harder benchmarks may have produced rising rates of fallacious items in newer benchmarks. Our findings have considerable implications for current benchmarking practice, and offer a scalable method for detecting fallacious items.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.