FedCURE: Mitigating Safety Erosion in Fully Benign Federated LLM Fine-Tuning through Cross-Client Update Consensus
Abstract
Fine-tuning aligned large language models (LLMs) on benign data can compromise safety even without malicious intent, but this phenomenon has been studied primarily in centralized settings. We show that safety erosion also arises in federated LLM fine-tuning, even when all participating clients are benign and train exclusively on benign local data. Our analysis reveals a federated-specific signal: erosion-aligned update directions that are consistently shared across clients can become disproportionately represented in the aggregated update. Motivated by this observation, we propose FedCURE, a safety-aware federated fine-tuning framework centered on Erosion Consensus Filtering (ECF). ECF identifies erosion-aligned candidates within individual client updates and filters those with strong cross-client consensus before aggregation. FedCURE further employs offline reference adapters to stabilize local adaptation and correct residual safety degradation in the accumulated global model. Experiments across diverse LLM families and downstream datasets demonstrate that FedCURE substantially reduces safety erosion while retaining effective task performance and general utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.