acceptodds
Under review as a conference paper at ICLR 2027

FluidMoE: An Asymmetric Scheduling Framework for Context-Parallel Mixture-of-Experts Training

Abstract

Mixture-of-Experts (MoE) training with Ulysses context parallelism (CP) and expert parallelism (EP) uses AllToAll (A2A) for attention and expert exchanges and AllReduce (AR) for gradient synchronization. Existing overlap methods largely optimize these paths separately. We introduce FluidMoE, an asymmetric scheduler that coordinates these exchanges and computation under a common critical-path communication exposure (CCE) objective. In forward, it decomposes each A2A into rounds of point-to-point exchange and overlaps data transfer with matrix multiplication. In backward, it schedules ready weight-gradient tasks at runtime to overlap A2A, then pipelines communication and activation-gradient chunks where exposure remains. It schedules AR on ready gradient slices within the available time between A2A exchanges, overlapping reduction with computation. Implemented as a scheduling layer in Megatron, FluidMoE reuses compute kernel implementations and preserves communication volume and training semantics. Across three public MoE families at 16K tokens, block-level speedups reach 1.2× against Megatron with gradient-reduce overlap enabled and 1.7× against DeepSpeed-MoE. Measured CCE, comprising A2A exposure and full tail-AR duration, falls by up to 67% relative to non-overlap Megatron. End-to-end loss trajectories remain aligned, and multi-node evaluation on data-center GPUs confirms the performance gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.