DBFB: Bit-Aware Critical Weight Monitoring for Defending Bit-Flip Backdoor Attacks
Abstract
Bit-flip backdoor attacks (BFBAs) constitute a sparse post-deployment threat in which an adversary flips only a few weight bits to activate targeted behavior while seeking to preserve clean accuracy. Existing bit-flip defenses are mainly designed for random faults, global accuracy degradation, or broadly sensitive parameters, and are therefore poorly aligned with the trigger-conditioned and highly localized nature of BFBAs. We propose DBFB, a selective defense that identifies and monitors BFBA-critical weights through a progressive neuron-to-weight-to-bit analysis. DBFB first locates attack-sensitive neurons by combining cross-class support with local boundary sensitivity and competition-direction concentration. It then maps these neurons to path-constrained candidate weights and evaluates their decision impact and structural propagation support. Finally, it re-ranks candidates using signed bit-induced changes and bounded estimates of relative margin erosion. The selected weights are stored with their clean reference values and checked before inference so that corrupted monitored weights can be detected and restored to their exact clean values. Across CNN and ViT architectures, four datasets, and three representative BFBAs, DBFB monitors only 5,000 weights, corresponding to 0.025% of model parameters on average. In the main evaluation, DBFB achieves 100% Hit Rate and 0% ASR for OneFlip on all four model–dataset pairs, and substantially suppresses the other evaluated attacks. The measured online check takes less than 0.1 ms, with approximately 39.1 KiB of index-and-reference storage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.