Rankcast: phase-aware rank balancing in MoE serving
Abstract
Modern large language model (LLM) inference increasingly adopts disaggregated architectures that separate Prefill (P) and Decode (D) phases. Meanwhile, Expert Parallelism (EP) has been widely adopted in large Mixture-of-Experts (MoE) models serving. However, load imbalance still slows down Mixture-of-Experts (MoE) inference because each layer must wait for the busiest EP rank. The main source of latency differs between prefill and decode. Prefill is compute-bound, and its MoE Layers' latency is determined by the maximum token count of the busiest rank. At the same time, under moderate batch sizes, Decode is memory-bound, so the latency is dominated by the maximum active-expert count across ranks. To solve both problems under a unified framework, we present RankCast, a system that reduces rank-level maximum token counts during prefill and active-expert counts during decode. For every token, RankCast preserves the predicted Top- experts and chooses the remaining from its Top- candidates (). In prefill, RankCast yields a price per rank that steers flexible assignments to less loaded ranks. In decode, RankCast closes experts on the busiest ranks while preserving the constraints for each token. Both paths are implemented and tested in the state of the art inference frameworks SGLang and DeepEP. We overlap the RankCast's kernel with model execution and reuses the existing DeepEP communication channel for meta data transmission. On Qwen3-235B-A22B with EP and Qwen3.5-397B-A17B with EP, RankCast removes over 75% of the busiest rank's excess prefill load and improves prefill input throughput by up to 19%, while reducing decode time per output token by up to 11%, nearly lossless.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.