Self-Correction with Integrity: From Trigger-Induced Gradient Alignment to Data-Preserving Backdoor Defense
Abstract
Poisoned samples are often learned earlier than clean samples, a phenomenon already exploited by backdoor defenses. However, how trigger-induced structure contributes to this learning advantage remains insufficiently understood. To this end, we propose Self-Correction with Integrity (SCI), which is a clean-set-free backdoor defense guided by a theoretical analysis of early learning. We decompose early training gradients to examine the roles of shared target labels and recurring trigger features. This analysis explains how these two factors can produce mutually reinforcing updates that favor early learning of the trigger-target association. Specifically, based on this analysis, SCI first uses low-capacity models to construct a suspicious subset with high poisoned-sample recall. It then uses a model trained on the remaining data to relabel suspicious samples according to their semantic content, weakening trigger-target associations. Next, SCI reconstructs the relabeled samples under the guidance of task-relevant features to suppress residual trigger information while preserving semantic content. Finally, the corrected samples are merged with the remaining data to form a purified training set that preserves data integrity for subsequent model training. Experiments on CIFAR-10 and ImageNet subsets show SCI achieves effective backdoor mitigation while maintaining clean accuracy. Comparisons with sample removal further demonstrate the accuracy benefit of retaining corrected samples. These results show that a mechanistic understanding of early learning can guide backdoor defense through sample correction, retaining the full training set without auxiliary clean data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.