Eraser: Erasing Trigger-Induced Effects for Backdoor Defense in Federated Learning
Abstract
Federated Learning (FL) is highly vulnerable to backdoor attacks, where malicious clients implant trigger-dependent behaviors into the global model while preserving main-task performance. Existing defenses rely on malicious-client identification or training-time intervention, and often suffer robustness degradation under heterogeneous data, stealthy attacks, and diverse federated settings. We propose Eraser, a post-training backdoor defense for FL that erases trigger-induced effects before inference, preventing triggers from activating backdoor behaviors in a poisoned global model. Since potential triggers are unknown and difficult to remove directly from input samples, Eraser instead reconstructs sanitized samples from robust latent semantics. Motivated by our observation that standalone benign local models exhibit inherent robustness against backdoor triggers, Eraser uses their Featurizers to suppress trigger-related cues in the latent space, while a collaboratively trained Generator reconstructs sanitized samples from these representations. To improve cross-client consistency, we introduce fixed hyperspherical semantic anchors to align local feature spaces and stabilize generator training. Eraser is model-agnostic, compatible with most federated algorithms, and requires neither trigger knowledge nor malicious-client identification. Experiments on diverse attacks and heterogeneous settings show that, compared with existing defenses, Eraser achieves the best robustness while preserving main-task accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.