Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling
Abstract
Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with expert networks, and each token is routed to of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On nodes of eight MI300X GPUs, these collectives can take 45 % of the training step at EP32 with top-2 routing and 60 % with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8 % of the expert pairs in a layer are selected together by 45 % of tokens, and the experts a token selects at one layer predict the experts it selects at the next layer. We use these correlations to keep more token–expert assignments on the token's own GPU, which reduces communication across GPUs and across nodes. Correlated expert placement puts experts that are often selected together on the same GPU. Combined with a dispatcher that sends each token to each GPU once, it removes up to 58 % of dispatched rows. Token shuffling applies when sequence parallelism shards tokens across the EP group. It moves each token to the GPU predicted to hold its next-layer experts during the reduce-scatter that follows attention. On one node this raises the share of token–expert assignments served on the token's GPU from 12.5 % to 59 %. In Megatron-LM, across EP degrees from 8 to 64 with top-2 and top-6 routing, the two methods reduce all-to-all time by 1.16–2.63 and end-to-end step time by up to 1.41. Neither method changes the models' underlying routing decisions or expert parameters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.