CLEANREF-LFDR: FDR-Controlled Contamination Detection with Clean References for LLM Evaluation
Abstract
Large-scale pretraining has enabled large language models (LLMs) to achieve strong performance across diverse tasks. However, overlap between evaluation benchmarks and pretraining corpora introduces test data contamination, undermining evaluation reliability by confounding memorization with generalization. Existing training data detectors seek to identify clean evaluation samples, but undisclosed training corpora make direct membership verification difficult. These detectors therefore rely on indirect signals and may leave contaminated samples in the retained set. False discovery rate (FDR) control provides a principled way to limit the expected proportion of contaminated samples retained while seeking to preserve as many clean samples as possible. Yet, approaches that rely on contaminated references for calibration face a practical challenge: when the training corpus is undisclosed, the membership of these references is difficult to verify. Trusted clean references can still be available for a fixed target model, such as private data known to have been excluded from training or genuinely new content created after the model’s release. This motivates us to pursue FDR control with high statistical power using only clean references. We propose CleanRef-LFDR, a sample selection framework that uses a finite trusted clean reference set and an unlabeled evaluation mixture. CleanRef-LFDR estimates the clean proportion and uses density ratios to infer local contamination risks, then selects the largest subset satisfying an average estimated-risk constraint. Accounting for proportion and density estimation errors, we establish a finite-sample FDR upper bound and asymptotic control under stated conditions. Experiments across models, datasets, and detection scores demonstrate effective empirical risk control and high statistical power in a broad range of controlled contamination settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.