ProtegoFed: Sample-Level Backdoor Defense for Federated Instruction Tuning with Interspersed Poisoned Data
Abstract
Federated instruction tuning (FIT) enables organizations to adapt large language models without sharing private instructions. Its training data, however, may be collected from untrusted users or third parties, allowing poisoned samples to be interspersed across otherwise benign clients. Existing federated backdoor defenses focus mainly on malicious clients and can be ineffective in this setting. We identify frequency-domain sample gradients as a signal for separating poisoned from clean instructions and introduce ProtegoFed, a one-time, three-stage defense that combines intra-client clustering, server-side secondary clustering of client centroids, and local revision using a global centroid. Across four free-style question answering datasets and four backdoor attacks, ProtegoFed identifies – of poisoned samples in the IID setting, reduces attack success rate to zero in all 16 dataset-attack combinations, and largely preserves main-task utility in those settings. Additional experiments cover nine models from 1B to 70B parameters, heterogeneous client distributions, clean-only data, design ablations, adaptive attacks, and integration with a client-level defense.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.