GARF: Gradient-Aligned Reactivation Forgetting for Backdoor Defense
Abstract
Backdoor defenses often suppress visible trigger behavior, yet leave latent backdoor features that can be reactivated by small, adversarially optimized perturbations. We study this problem in a realistic post-training setting where the defender only has access to a possibly backdoored model, without any poisoned data. We propose Gradient-Aligned Reactivation Forgetting (GARF), a bilevel post-training defense that treats the parameter difference between the original backdoored model and its defended counterpart as a proxy backdoor task direction in parameter space. In the lower level, GARF learns a small input-agnostic perturbation on clean data whose induced parameter gradients are encouraged to align with this direction, thereby simulating an adaptive reactivation attack using only accessible information. In the upper level, the model is updated to jointly reduce the clean loss and a reactivation loss evaluated under the learned perturbation, forming an iterative attack–defense process that repeatedly exposes and suppresses reactivation-sensitive directions. Empirically, GARF reduces the white-box reactivation attack success rate by over 50% on average and keeps the black-box reactivation ASR below 10%, while preserving competitive clean accuracy across multiple benchmark datasets and diverse attack types.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.