SCAR: Class-Wise Assessment and Shortcut-Guided Model Revision for Backdoor Defense
Abstract
Backdoor attacks pose a severe security threat to deep neural networks (DNNs) by implanting hidden behaviors activated by specific triggers. This threat is particularly concerning for downstream users acquiring models from untrusted sources, for whom reliable deployment entails assessing model integrity and, if compromised, localizing suspicious target classes and repairing the model. Backdoor learning introduces an additional trigger-to-target association beyond the intended task mapping, forming a shortcut that can leave class-specific irregularities in classifier-head weights and prediction responses, as well as directional changes in feature space. Building on these observations, we propose Shortcut-aware Class-wise Assessment and Revision (SCAR), a unified post-training framework for backdoor assessment and model recovery. SCAR first combines classifier-weight anomalies with perturbation-induced prediction preferences to assess model integrity and localize suspicious target classes. For each localized class, it then estimates a shortcut-related feature direction from class-specific weights and feature-wise variability, and selectively revises the corresponding classifier-head parameters along this direction to suppress backdoor behavior while preserving normal predictive utility. Extensive experiments across diverse datasets, attacks, and architectures demonstrate the effectiveness and generality of SCAR for both backdoor assessment and model recovery. Code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.