acceptodds
Under review as a conference paper at ICLR 2027

BlindClean: A Prior-Knowledge-Free Data Sanitization Method against Clean-Label Backdoor Attacks

Abstract

Machine learning training commonly relies on third-party datasets to reduce the costs of data collection and annotation, but this practice also exposes models to clean-label backdoor attacks. Such attacks are difficult to detect, as poisoned samples retain both their original labels and their visual semantics. Existing data sanitization methods either depend on prior knowledge unavailable to the platform or degrade clean accuracy by mistakenly discarding too many clean samples. This paper proposes BlindClean, which leverages a key characteristic of poisoned samples: under input perturbations, poisoned samples tend to exhibit higher feature stability than clean samples. BlindClean therefore sanitizes a dataset relying solely on the potentially poisoned dataset itself. It evaluates cross-layer feature stability under multi-scale input perturbations, estimates reference feature distributions of clean samples from low-stability samples of each class, and then measures the feature deviations of high-stability samples from these distributions. This design allows BlindClean to first identify the class most likely to be attacked and then detect the poisoned samples within it. Experiments on three representative clean-label backdoor attacks show that BlindClean detects at least 97% of the poisoned samples while mistakenly removing less than 1% of clean samples. When the sanitized data is used for downstream training, BlindClean reduces the average attack success rate to below 1% and improves clean accuracy by up to 44 percentage points over state-of-the-art methods. BlindClean thus provides a practical and effective means of sanitizing untrusted third-party data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.