RouteEP: Routing-Aware Expert Parallelism for Efficient MoE LLM Decoding
Abstract
Mixture-of-Experts (MoE) architectures enable Large Language Models (LLMs) to scale their parameter capacity while activating only a small subset of experts for each token. However, dynamic expert routing in expert parallelism (EP) and expensive all-to-all collective communication incur substantial inter-device communication overhead during model inference. Such bottlenecks become more pronounced in decoding, where sparse expert activation at small active batch sizes limits computation–communication overlap and effective parallelism. In this paper, we present RouteEP, a routing-aware expert parallelism framework to reduce inter-device communication and improve the load balance of expert execution for LLM decoding. RouteEP introduces a decode-specific parallelism design tailored to the memory- and communication-bound nature of decoding through three coordinated contributions. (1) We propose a lead-anchored offline expert grouping algorithm that selects lead experts, builds expert groups around them, and assigns the remaining experts to maximize position-wise routing affinity under capacity constraints. (2) Based on these expert groups, we design a routing-aware expert parallelism scheme for LLM decoding that replaces the communication-intensive dispatch–combine EP pattern with rank-local partial expert computations followed by a global all-reduce. (3) We further integrate intra-expert sharding into this framework to increase parallelism within the limited set of activated experts. Experiments on a range of MoE LLMs, including Qwen, DeepSeek, and Mixtral models, show that RouteEP achieves – speedups over state-of-the-art inference frameworks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.