FedPDPO: Federated Personalized Direct Preference Optimization
Abstract
Direct Preference Optimization (DPO) has emerged as an efficient alternative to reinforcement learning from human feedback (RLHF) for aligning large language models (LLMs), eliminating the need for explicit reward models and online policy optimization. However, applying DPO in federated settings introduces a fundamental challenge: different clients may exhibit heterogeneous preference distributions, making a single implicit reward function insufficient for personalized alignment. In this work, we identify that federated preference alignment requires jointly preserving shared linguistic knowledge and adapting to client-specific preference signals. To address this challenge, we propose FedPDPO (Federated Personalized Direct Preference Optimization), a personalized federated framework that decouples global representation learning from local preference calibration. Specifically, FedPDPO employs a LoRA-adapted global backbone shared across clients and personalized local preference heads that capture client-specific reward variations without sharing private preference data. Furthermore, we introduce PDPO, a personalized preference optimization objective that augments DPO's implicit reward with explicit client-specific reward corrections, providing a client-specific pathway for modeling preference differences. We provide a probabilistic interpretation of the reward correction mechanism based on the Bradley–Terry preference model. Extensive experiments on diverse preference alignment benchmarks demonstrate that FedPDPO consistently outperforms existing federated alignment and personalized FL baselines, on the reported held-out preference-pair accuracy across the evaluated client partitions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.