Guarding Federated LLM Finetuning: Filtering Unsafe Clients That Break Alignment
Abstract
Federated finetuning of LLMs is a promising solution to collaboratively specialize an LLM for domain specific tasks without sharing private data. However, most prior works do not consider the risk of malicious clients attempting to jailbreak the safety of LLM. Traditional malicious client detection approaches in federated learning typically rely on statistical outlier detection using client updates. These approaches are less effective in federated LLM finetuning where benign clients can be highly heterogeneous due to real-world data distribution differences across clients and the ability of LLMs to be finetuned for diverse tasks within a unified generative framework. Existing safety-oriented defenses for federated LLMs also often rely on an impractical assumption about knowledge of client dataset distribution or introduce unfavorable trade-offs between improvement in safety and degradation in task performance. To address this, we propose a safety-aware method for detecting unsafe clients rather than relying solely on raw statistical deviations between client updates. Specifically, we evaluate the displacement in weights brought by each client relative to a reference safety degrading parameter movement defined between a safety-aligned and an unsafe model. By relying on a publicly available harmful dataset alone to obtain the unsafe model, we make no assumption about knowledge of client's private data. Experiments across different poisoning, heterogeneity levels and architectures show that our approach can effectively distinguish unsafe clients and obtain the best safety-utility tradeoff after federated finetuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.