MoEAtlas: Orchestrating Communication in MoE Inference via Pre-Routing, Expert Placement, and Migration
Abstract
DeepSeek reports substantial communication overhead in mixture-of-experts (MoE) inference with expert parallelism across nodes. Each MoE layer sends tokens to GPUs holding their selected experts. Placing experts selected together on fewer nodes reduces transfers; choosing where each request runs makes more transfers local. We present MoEAtlas, which coordinates Pre-Router, hypergraph expert placement, and online migration to reduce communication in distributed MoE inference. The offline-trained Pre-Router predicts expert selections of prompt tokens across MoE layers and uses these predictions with the current expert layout to select a home GPU for each request. We represent each token's co-activated experts as a hyperedge and use hypergraph partitioning algorithms to produce new expert placements to reduce inter-node communication under load and GPU memory budgets. As workloads change, our migration strategy optimizes the mapping of the target expert placement to devices to reduce migration volume and time while inference continues. New hypergraph expert layouts with Pre-Router reduce complete prefill time by 19.33%, 17.00%, and 24.69% on OLMoE, Qwen1.5-MoE, and DeepSeek-V2-Lite, respectively, compared with the original layouts. On OLMoE, min-volume mapping reduces inter-node weight-transfer volume by 11.95% relative to fixed mapping; online migration reduces the total time to serve requests and complete a layout update by 6.83% relative to blocking migration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.