acceptodds
Under review as a conference paper at ICLR 2027

Fighting Fire with Fire: Assessing Test Set Contamination via Deliberate Training on Test Data

Abstract

Test set contamination inflates benchmark scores and misrepresents what a model can do. Existing detectors read a trained model’s static outputs, which a developer can optimize against directly. We audit by intervention instead: fine-tune the suspect model on the benchmark in question and read the shape of its loss curve. A model re-exposed to data it has seen before descends differently than one meeting that data for the first time, and the difference survives zero-normalizing every curve, so it lives in the shape of learning rather than in the loss level. We introduce FIRE (Fine-tuning-Induced Re-Exposure), which classifies these curves with dynamic time warping and k-nearest neighbors: no learned parameters, no model internals, one fine-tuning run (about ten minutes for a 7B model on a single GPU) and a lookup against a reference library built once per benchmark. Across six models (1B–7B), 44 benchmarks, and over 43,000 training runs, FIRE reaches 77–92% accuracy at contamination levels that meaningfully inflate benchmark scores, matching or beating five membership inference baselines and beating every baseline that does not require a clean same-architecture model. The signal transfers across model families at 87%, reaches a 70B model at 69–78% zero-shot and 88.7% with 50 labeled trajectories from it, and persists through a subsequent preference-optimization (DPO) stage with clean data. Detection needs as few as 256 benchmark samples, while evading it requires a full probing run inside every poisoning step. What a model has already seen is legible in how it learns.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.