acceptodds
Under review as a conference paper at ICLR 2027

DIFD: Pre-Training Poisoned Sample Screening via Timestep-Calibrated Diffusion Residuals

Abstract

Backdoor attacks embed covert triggers into training data to induce attacker-specified results during inference. To prevent these malicious shortcuts from being encoded into model parameters, screening poisoned samples prior to model training is essential. However, pre-training screening faces two key hurdles: defenders lack access to a trained classifier or trigger priors, and backdoor triggers exhibit vast heterogeneity across spatial and spectral domains, making fixed-rule detectors hard to generalize. To address these challenges, we propose DIFD, a pre-training poisoned sample screening framework via timestep-calibrated diffusion residuals. We observe that intermediate reconstruction residuals of poisoned and benign samples exhibit timestep-dependent separability, while inadequately reconstructed trigger components remain in the residual as spatial structural or spectral anomalies. Based on these observations, DIFD first performs timestep calibration to locate an intermediate reconstruction window where this separability is maximized. It then employs spatial-spectral residual scoring to jointly measure spatial structural deviations and spectral distribution shifts, producing a unified sample-level risk score. DIFD relies solely on a fixed diffusion model pretrained on clean data, requires no auxiliary detector training or access to a downstream classifier, and avoids full reverse diffusion sampling. Experiments across multiple datasets, image resolutions, and attack settings demonstrate that DIFD achieves stable cross-attack detection performance and efficient poisoned sample screening.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.