FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data
Abstract
Recent advances in language models have established reinforcement learning as the primary paradigm for eliciting self-correction and long-chain reasoning. While group relative policy optimization (GRPO) offers superior scalability by eliminating the critic network, deploying it on a central infrastructure entails collecting a large volume of data from distributed owners, which poses communication overhead and significant privacy risks. To address these concerns, we introduce federated GRPO (FGRPO), a framework designed to decentralize the fine-tuning of reasoning models across different data owners. To account for differences in local learning progress under heterogeneous client data, FGRPO incorporates an adaptive aggregation mechanism based on relative performance gain. By characterizing each client's improvement relative to its historical baseline, the framework dynamically prioritizes effective learning trajectories regardless of local data heterogeneity. Rigorous theoretical analysis and extensive experiments are conducted to verify the efficacy of FGRPO.Our code is available at https://anonymous.4open.science/r/FGRPO-9BF8.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.