TIDE: Defending Against Backdoor Data Poisoning via Dense Influence Subgraph Detection
Abstract
Targeted poisoning attacks perturb the training data of machine learning models to increase the probability that the trained models make certain errors. Stealthy backdoor attacks, which perturb data by injecting inconspicuous triggers into the features, pose a particularly serious threat to the reliability and security of machine learning. To detect and filter backdoor-poisoned data, we shift the focus from identifying individually anomalous data points to identifying groups of training data that have a disproportionately high impact on each other during training. Specifically, we first introduce an efficient approach for estimating how the inclusion of a data point in model training influences the output of the trained model for another data point. Leveraging the observation that this influence is disproportionally high between pairs of poisoned data, we introduce an efficient algorithm for detecting poisoned data as a dense subgraph in the graph formed by data points and the influence values between them. We evaluate our proposed approach against state-of-the-art backdoor poisoning attacks on benchmark datasets, demonstrating that it outperforms existing approaches for data cleaning in terms of reducing attack success and increasing model accuracy. We also demonstrate robustness against an adaptive attack that attempts to circumvent our defense by reducing the cross-influence of poisoned data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.