acceptodds
Under review as a conference paper at ICLR 2027

Influence-Guided Active Search for Poisoned Training Data Forensics

Abstract

Data poisoning attacks that inject malicious samples into training data pose a serious threat to the reliability of machine learning. Existing defenses focus on automated detection and removal, but effective cleaning inevitably discards significant amounts of benign data. We instead consider forensic investigations of poisoned data, which verify samples through manual inspection or provenance checks, such as re-acquiring a sample from its original source or auditing its origin. A key challenge of forensic investigations is that verification is expensive and can be performed only per sample, but investigation budgets are limited. Therefore, an investigation must strategically select the samples that will be verified and potentially removed to minimize the impact of the remaining poisons. We frame this as a nonmyopic sequential search problem and introduce an influence-guided active search (IGAS) framework that integrates: (i) a label-free influence score identifying samples with disproportionate impact on predictions and (ii) an adaptive search strategy propagating information from verified samples to focus on regions of the data that are both influential and likely to be poisoned. Experiments on CIFAR-10, Tiny ImageNet, and GTSRB across eight triggerless and backdoor attacks demonstrate that IGAS substantially reduces Attack Success Rate while preserving test accuracy, outperforming automated data-cleaning defenses that often remove benign samples along with poisoned ones.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.