LigerMoE: Efficient Expert Parallelism via Warp-Specialized Communication-Compute Fusion
Abstract
State-of-the-art large language models rely on Mixture-of-Experts (MoE) layers to scale model capacity while maintaining computational efficiency. At scale, MoE models rely on expert parallelism, whose efficiency is limited by many small expert matrix multiplications and expensive all-to-all communication. To address these bottlenecks, we present persistent, fused MoE kernels for both the forward and backward passes that leverage warp specialization to overlap communication with computation. The kernels use remote direct memory access (RDMA) to access data on remote GPUs without bidirectional sender-receiver coordination, while preallocated symmetric buffers support CUDA Graph execution without requiring a predefined per-expert token capacity. On the Qwen3-30B-A3B MoE model, our forward kernel achieves and higher throughput than the strongest available baselines on H200 and B300 GPUs, respectively. For the backward pass, our method improves throughput by and on the same systems. In end-to-end BF16 training on H200 GPUs, our implementation reduces training step time by compared with Megatron's default MoE implementation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.