From Detection to Intervention: Counterfactual Decontamination for Language-Model Evaluation
Abstract
Benchmark decontamination is commonly treated as an overlap-detection problem, even though evaluation signal can enter training data through paraphrases, answer-bearing context, and reusable solution templates. We introduce Counterfactual Decontamination, an intervention-first framework that represents benchmark-to-training links as typed edges and maps them to auditable data actions. The framework combines lexical, semantic, answer, and structural evidence, then constructs matched contaminated, exact-pruned, graph-pruned, reweighted, and oracle-pruned training arms. This design separates three questions that are often conflated: whether a candidate edge can be retrieved, whether it constitutes actionable leakage, and whether removing it changes the evaluation conclusion. In controlled real-text mathematics contamination, exact matching recovers copies but misses non-exact channels, whereas the typed graph recovers paraphrase, answer, and template exposures while retaining hard benign neighbors. Across contamination levels and random seeds, exact pruning reduces but does not eliminate the target-to-clean training-effect gap; graph pruning closes 99.1% of the gap and approaches the oracle intervention under matched evaluation. Natural-pool blind validation and matched clean controls further test whether high-severity edges remain actionable beyond controlled injection. The resulting protocol replaces overlap counts with a falsifiable criterion: a decontamination policy is supported only when its detected edges survive blinded validation and its intervention preserves the model comparison the benchmark is meant to measure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.