acceptodds
Under review as a conference paper at ICLR 2027

FedSHIELD: Adversarial Honeypot-Guided Behavioral Client Screening for Robust Federated Learning

Abstract

Federated Learning (FL) enables collaborative training of the model between distributed clients without sharing raw data, providing significant privacy benefits. However, such training is highly vulnerable to backdoor attacks, where malicious clients insert hidden noise that compromises the global model while preserving apparent performance on clean data. Existing defenses struggle under realistic FL conditions, where benign updates exhibit high variability, and adaptive attackers can avoid statistical detection. To address these challenges, we propose FedSHIELD, a server side, honeypot guided behavioral client screening framework which evaluates client updates by measuring their reduction in loss on adversarially crafted honeypot samples. This approach enables direct detection of malicious updates without relying on layer wise statistics or prior assumptions about attack patterns. Suspicious updates are filtered out, while the remaining updates are clipped and perturbed with differential privacy noise to ensure robustness and privacy preservation. Extensive experiments on benchmark datasets such as MNIST, FashionMNIST, SVHN, and CIFAR-10 demonstrate that FedSHIELD achieves substantially lower backdoor attack success rates while maintaining high clean-data accuracy. FedSHIELD outperforms existing defenses across both IID and non-IID data distributions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.