Collaborative Specialist Knowledge Distillation for Federated Offline Reinforcement Learning
Abstract
Federated offline reinforcement learning coordinates policies trained from decentralized offline experience, yet model-level aggregation can underutilize complementary behaviors acquired by individual clients. We propose Collaborative Specialist Knowledge Distillation (CSKD), which preserves client-specific specialization across communication rounds and transfers specialist behavior into a unified shared policy. Each client maintains a persistent residual adaptation and performs local Conservative Q-Learning, while the server composes state-policy targets contributed by multiple specialists and integrates them through policy-space distillation. On heterogeneous D4RL MuJoCo benchmarks, CSKD achieves normalized scores of 48.04 on HalfCheetah, 100.16 on Hopper, and 88.54 on Walker2d, with an average score of 78.91. Additional experiments examine repeated collaboration, client scalability, communication cost, and experience overlap, showing that specialist knowledge can be transferred effectively with substantially lower client-to-server payload than full-model exchange.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.