Fighting Fire with Fire: Assessing Test Set Contamination via Deliberate Training on Test Data
Abstract
Test set contamination inflates benchmark scores and misrepresents what a model can do. Existing detectors read a trained model’s static outputs, which a developer can optimize against directly. We audit by intervention instead: fine-tune the suspect model on the benchmark in question and read the shape of its loss curve. A model re-exposed to data it has seen before descends differently than one meeting that data for the first time, and the difference survives zero-normalizing every curve, so it lives in the shape of learning rather than in the loss level. We introduce FIRE (Fine-tuning-Induced Re-Exposure), which classifies these curves with dynamic time warping and k-nearest neighbors: no learned parameters, no model internals, one fine-tuning run (about ten minutes for a 7B model on a single GPU) and a lookup against a reference library built once per benchmark. Across six models (1B–7B), 44 benchmarks, and over 43,000 training runs, FIRE reaches 77–92% accuracy at contamination levels that meaningfully inflate benchmark scores, matching or beating five membership inference baselines and beating every baseline that does not require a clean same-architecture model. The signal transfers across model families at 87%, reaches a 70B model at 69–78% zero-shot and 88.7% with 50 labeled trajectories from it, and persists through a subsequent preference-optimization (DPO) stage with clean data. Detection needs as few as 256 benchmark samples, while evading it requires a full probing run inside every poisoning step. What a model has already seen is legible in how it learns.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.