Scalable and Decentralized Training of Mixture of Experts with Collaborative Gating
Abstract
Decentralized federated learning with Mixture-of-Experts (MoE) allows clients to directly share lightweight experts and combine them locally, avoiding the server dependency and knowledge degradation of centralized aggregation. However, existing decentralized FL-MoE methods assume fully-connected communication among all clients, which incurs quadratic per-round cost and does not scale. In addition, each client’s gate is trained only on local data, which is insufficient to learn reliable routing under non-IID distributions. We present DCGMoE (Decentralized Collaborative Gating MoE), a decentralized FL-MoE framework that addresses both issues. First, DCGMoE introduces a data-aware grouping strategy that partitions federations into groups of complementary clients using local statistics such as label distribution. Each client then communicates only with members of its own group, which reduces the per-round communication cost from quadratic to sub-quadratic growth as the federation size increases. Second, DCGMoE proposes a collaborative gating mechanism that jointly trains the routing module across clients so that each gate learns to exploit the expert pool within its group. The proposed DCGMoE achieves higher global accuracy than both centralized and decentralized FL baselines on Fashion-MNIST, CIFAR-10, and CIFAR-100 under non-IID partitioning with α=0.3. At a federation of 100 clients, DCGMoE further achieves substantially lower per-round communication cost and shorter per-round training time than baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.