MoE-DisCo: Low Economic Cost Training Mixture-of-Experts Models
Abstract
Training large-scale Mixture-of-Experts (MoE) models faces prohibitive memory capacity constraints. While only a subset of parameters is activated during inference, the entire parameter space must be updated during training. This forces the use of multi-GPU parallelism (e.g., Tensor Parallelism), introducing massive inter-GPU communication overhead that significantly bottlenecks training efficiency. To address this, we propose MoE-DisCo (Mixture-of-Experts with Disentangled Clustering and Coordination)—a staged, communication-efficient training framework. MoE-DisCo decomposes the massive MoE model into multiple dense submodels, each consisting of a shared backbone and a single expert, ensuring each submodel fits within the memory of a single GPU. Concurrently, it partitions the training data into subsets using unsupervised clustering. Each submodel is trained independently and in parallel on its assigned data subset on isolated GPUs, completely eliminating inter-GPU communication overhead. Subsequently, all experts are integrated into a complete MoE model and fine-tuned globally for a short period. Experiments demonstrate that MoE-DisCo matches or exceeds the performance of full-parameter training across downstream tasks and perplexity (PPL), while drastically reducing training cost and communication overhead on architectures like Qwen-MoE and Llama-MoE.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.