DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
Abstract
Mixture-of-experts (MoE) architectures enable trillion-parameter LLMs with sparsely activated experts. Expert parallelism (EP) is a widely adopted MoE training strategy, but it suffers from severe all-to-all communication bottlenecks, exacerbated by limited inter-node bandwidth as growing model sizes require distributing experts across GPU nodes. Prior work focuses on overlapping these communications with feed-forward network (FFN) and self-attention computations, which often leaves residual network-bound stalls due to the imbalance in attention and FFN layers' computation–communication ratios. We present DisagMoE, a disaggregated MoE training system that jointly optimizes model placement and scheduling for efficient training. DisagMoE separates attention and FFN layers into disjoint GPU groups, introduces AF-Pipe, a multi-stage pipeline with a unidirectional forward path and many-to-many communication, and employs a network–compute roofline model to guide GPU and network resource allocation between the groups. Implemented on Megatron-LM, DisagMoE achieves up to 1.80x throughput over Megatron-1F1B across multiple MoE models in controlled cross-node EP experiments on 16-node clusters with eight H800 GPUs per node.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.