APTI-FL: Adversarial Perturbation Guided Trigger Inversion for Backdoor Detection in Federated Learning
Abstract
Federated learning limits the server's visibility into local training, allowing compromised clients to inject backdoors into a shared model. Existing approaches struggle to detect such malicious clients when non-IID data makes benign updates closely resemble malicious ones, or when attackers constitute the majority of participating clients. In our preliminary study, we discover a distinctive property of backdoored models: although adversarial perturbations substantially disrupt normal classification paths, a triggered input remains likely to be classified into the specified target even after the same perturbation is added. In other words, the backdoor path exhibits target retention under adversarial perturbations. Motivated by this observation, we propose APTI-FL, an adversarial attack guided federated backdoor defense. APTI-FL first generates untargeted adversarial perturbations from trusted server data and uses discrete and continuous region inversion paths to recover potential triggers with different spatial structures. It then determines a reliable attack target through multistage validation of candidate trigger paths and target correction across rounds. Finally, it combines multidimensional client behavior with historical evidence to identify clients that exhibit backdoor responses and blocks their contributions before aggregation. Experiments on five benchmark datasets against five representative defenses show that APTI-FL substantially reduces attack success rates while maintaining main task accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.