acceptodds
Under review as a conference paper at ICLR 2027

Benchmark Deja Vu: Debiasing Signals for Multiple-Choice Contamination

Abstract

Benchmark contamination compromises the reliability of evaluations used by frontier labs to measure progress and by users to compare language models. Because ordinary model behavior can mimic contamination signals, reliable detection remains challenging. We introduce three contamination detectors with explicit null hypotheses that adjust for specified response preferences. We apply these new detectors to 35 contemporary open and closed models across four widely used multiple-choice benchmarks: MMLU, MMLU-Pro, GPQA Diamond, and the multiple-choice subset of Humanity’s Last Exam (HLE). This audit reveals that most models produce at least one statistically significant signal on MMLU, whereas signals on other benchmarks are more selective, with family members sometimes sharing signals on the same benchmark. We aggregate detector-specific e-values to rank models by the strength of evidence against the global null across detectors and benchmarks. In addition, we construct two new testbed datasets, WikiCutoffQA and SyntheticSourcesQA, whose questions are screened so that model accuracy before exposure is consistent with random guessing. With these datasets, we can induce controlled exposure by training models on the original questions (or on versions with shuffled choices or paraphrased wording). Using these testbeds, we compare our detectors with 19 published contamination detection methods and demonstrate leading detection performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.