Beyond Single-Policy Federated LLM Alignment under Conflicting Preferences
Abstract
Federated LLM alignment seeks to align models with human values without sharing private data, while reconciling global consensus with diverse user preferences. Recent online iterative methods, which fine-tune models on freshly sampled responses, have shown strong promise for improving alignment quality. We present FedDePO, a framework that brings online preference optimization into federated learning while mitigating the performance degradation caused by conflicting preferences. To ensure reliable local supervision, FedDePO introduces a preference transitivity metric that filters noisy labels by measuring the logical self-consistency of judgments. Leveraging the implicit reward structure of the optimization, it further detects cross-client preference divergence without exposing raw data. To handle preference heterogeneity, FedDePO adopts a dual low-rank adapter design that structurally decouples consensus learning from personalized adaptation, and a policy-timestamp alignment mechanism that adaptively schedules local steps to preserve data freshness under communication constraints. Experiments across multiple model scales show that FedDePO achieves state-of-the-art results, closely approaching centralized online training while maintaining strict privacy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.