acceptodds
Under review as a conference paper at ICLR 2027

Why Can Policy Distillation Evade Data Contamination Detection?

Abstract

Data contamination detection is essential for reliable LLM benchmark evaluation. Conventional detectors often exploit unusually confident predictions associated with overfitting to training data. However, we find that policy distillation can expose students to evaluation data without producing this signature, as soft-target training encourages teacher agreement rather than certainty about corpus tokens. We theoretically show that distillation exposure can remain undetected by conventional detectors, whereas hard-label training can become detectable. To systematically study this problem, we construct Distillation-MIA, a benchmark for contamination detection in distillation post-training. Experiments show that conventional detectors perform strongly under matched supervised fine-tuning but fall close to chance under both online and offline distillation. We further propose two teacher–student KL-based detection methods, PD-Div and PD-Div-Gain. PD-Div-Gain improves average AUROC over the strongest conventional detector by 13.1% and 12.4% under online and offline distillation. Finally, an exploratory Qwen3 audit extends our analysis to a public-model setting. Our findings reveal the difficulty of detecting contamination introduced by policy distillation, which remains an important problem for further investigation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.