What Does “No Contamination Detected” Certify? Dose-Calibrated Detection Limits for Benchmark Contamination
Abstract
Benchmark contamination threatens the reliability of large language model evaluation, yet a negative contamination audit is difficult to interpret without knowing the sensitivity of the detector being used. Existing audits often report “no evidence of contamination” without quantifying what amount of exposure would have been detectable. We introduce CONTAM-LADDER, a controlled framework for calibrating contamination detection sensitivity. CONTAM-LADDER constructs models with known contamination exposure levels through dose-controlled training and evaluates detectors through power curves rather than binary decisions. This enables estimating minimum detectable doses (MDDs) and determining when a negative audit result is informative versus inconclusive. Across 36 contamination cells spanning model scales, benchmarks, exposure depths, and surface forms, we reveal a persistent gap between evaluation impact and detector sensitivity. In particular, at 1.4B parameters, training on four verbatim copies of BBH improves benchmark accuracy by 8.6 points across two seeds, while the strongest evaluated likelihood-based detector achieves only 0.39 per-item detection power. We further show that this silent region persists under isolated-dose validation and separate benchmark-specific memorization effects from broader answer-format transfer. Our findings demonstrate that contamination audits should report calibrated sensitivity rather than only negative findings. CONTAM-LADDER provides a reproducible framework for measuring what contamination detectors can and cannot rule out.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.