ChunkBalance: Chunk Redistribution for Data-Parallel Balancing in MoE Prefill
Abstract
Mixture-of-Experts (MoE) models are typically deployed with data parallelism (DP) and expert parallelism (EP). In real-world MoE prefill scenarios, where requests are long and exhibit substantial variation in arrival rates and sequence lengths, DP-EP deployment can suffer from severe DP imbalance caused by two synchronous all-to-all communication barriers, leading to degraded system performance. To address this problem, we propose ChunkBalance, a data-parallel balancing method based on chunk redistribution for MoE prefill. ChunkBalance is built on two-chunk splitting and compute-communication overlap. We observe that: (1) unlike EP workloads, which are bound to specific ranks, DP workloads can be flexibly redistributed across the DP group; and (2) synchronization occurs only at the boundaries between the DP and EP stages, while there is no synchronization barrier within the DP stage. Based on these observations, ChunkBalance redistributes the DP workloads of the second chunk across the DP group, allowing workloads across chunks to compensate for one another and thereby improving load balance. ChunkBalance operates offline and can be combined with request-level scheduling algorithms and heterogeneous parallelism configurations. Its communication overhead is nearly 2/3 smaller than the total intermediate-state size, while the remaining overhead can be effectively hidden through self-overlapping. Experiments across various workload and system configurations show that ChunkBalance reduces MoE prefill end-to-end latency by up to 40%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.