Tracing How Backdoors Are Learned: Poisoned Sample Detection via Information-Trajectory Distortion
Abstract
Deep neural networks (DNNs) are vulnerable to backdoor attacks, in which adversaries inject training samples containing specific trigger patterns to induce a malicious association between the trigger and a target label. We find that such attacks leave detectable traces during training. Specifically, within the target class, clean and poisoned samples are learned through different predictive cues, primarily normal semantics and trigger patterns, respectively, causing their information dependencies to progressively diverge. Meanwhile, poisoned samples exhibit coherent changes because they share the same attack target, further inducing an overall class-level shift. In contrast, benign classes exhibit no such intra-class learning divergence and remain relatively stable. We define this phenomenon as Class-Conditional Information-Trajectory Distortion (CITD). Based on this observation, we propose PGIT, a training-set purification method guided by information-learning trajectories. To characterize CITD, we use mutual information to capture the dependencies among the input, internal representations, and predictions, and model their evolution throughout training as information-learning trajectories. We then jointly measure the distance between sample trajectories and their clean class references and the displacement of the class-mean trajectory to localize potential target classes. Within each candidate class, we partition samples according to directional differences in temporal activation gradients and use trusted reference gradients to identify the benign partition, thereby detecting poisoned samples. Extensive experiments show that PGIT consistently outperforms existing baselines, achieving 100.00% TPR and 0.03% FPR on CIFAR-10/ResNet-18 under the LC attack. On fully clean benchmark datasets, PGIT achieves 0.00% FPR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.