Residual-Direction Filtering for Backdoor Defense in Federated Learning
Abstract
Malicious clients in federated learning can implant backdoors through poisoned updates, and server-side defenses must filter such updates without clean data or knowledge of the attackers. Detection-based defenses increasingly score clients on parameter subsets chosen by dedicated mechanisms, such as importance ranking or critical-layer selection, yet it is unknown when scoring a subset yields the same filtering decisions as scoring all parameters. We propose CoPR, which scores each client on a fixed parameter subset by the direction of its cross-client mean-centered residual relative to the coordinate-wise median residual, filters clients with a median-absolute-deviation threshold, and aggregates the clipped full updates of retained clients. We give sufficient conditions under which residual directions separate malicious from benign clients, show that a parameter subset preserves every filtering decision whose margin exceeds the subset-induced score perturbation, and bound the single-round aggregation error under missed detections, false rejections, and clipping. Consistent with this criterion, subsets holding only 0.78% of the parameters match full observation on federated diffusion models at every tested location with 13–54× lower detection time, whereas non-IID classification requires full observation. Across five attacks on CIFAR-10 diffusion models, CoPR attains the lowest average FID while suppressing private-image reconstruction as well as AlignIns and remains effective against a subset-aware variant of AdaSCP; in non-IID classification, CoPR-full reduces the average attack success rate to 6.86% at 82.78% clean accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.