FedOPD: Private On-Policy Distillation of Heterogeneous Client Knowledge into Foundation Models
Abstract
Transferring decentralized domain expertise to a centralized foundation model is critical for federated foundation models (FedFMs). However, existing methods either expose model states via parameter sharing or risk training data leakage through client-side content generation and outcome-reliant feedback. We propose **FedOPD**, an on-policy distillation framework that renders clients entirely *generation-free* and *outcome-free*. In FedOPD, the server dispatches its own reasoning rollouts to domain clients, whose frozen experts passively evaluate the received tokens via teacher forcing to return dense, token-level log-probabilities. This discrepancy directly optimizes the sequence-level reverse KL divergence with an variance bound, bypassing coarse verifiers and memorization leakage. To facilitate practical deployment, we develop **SPBO** to prevent catastrophic forgetting across sequential domains and **TTF** to prune feedback communication by up to 80%. Extensive experiments on mathematical and sequential multi-domain benchmarks show that FedOPD outperforms prior federated baselines and centralized verifier-guided GRPO across 3B and 7B students, while using only 1% of FedDF's communication bandwidth.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.