acceptodds
Under review as a conference paper at ICLR 2027

Distribution Aligned Reweighting for Backdoor Defense

Abstract

Backdoor attacks pose a serious threat to deep neural networks by embedding a backdoor to induce malicious behavior on triggered inputs while preserving normal performance on clean inputs. Existing in-training defenses often rely on assumptions about differences in the learning dynamics of clean and poisoned samples, which may not hold consistently across attacks. We propose Distribution-Aligned Reweighting (DAR), an in-training backdoor defense that views backdoor poisoning as an adversarial shift of the training distribution away from the underlying clean distribution. Given a small trusted clean reference set, DAR performs class-wise distribution alignment to assign weights to training samples and uses the resulting weights to construct a sampling distribution for standard model training. DAR does not require prior knowledge of the attack and does not modify the model architecture or training objective. Extensive experiments across CIFAR-10, GTSRB, and ImageNet, covering diverse backdoor attacks and poisoning ratios, demonstrate that DAR consistently achieves low attack success rates while maintaining high clean accuracy and strong robust accuracy on triggered inputs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.