Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Execution
Abstract
Mixture of Experts (MoE) architecture has become the standard for state-of-the-art large language models, owing to its computational efficiency through sparse expert activation. However, sparsity through finer expert granularity is becoming increasingly difficult to achieve due to fundamental training challenges such as expert collapse and load imbalance. In this work, we explore and leverage intra-expert activation sparsity as a complementary and underexplored dimension of sparsity in MoE models. Surprisingly, substantial intra-expert sparsity is readily available in off-the-shelf pre-trained MoE models, providing up to 98% sparsity on top of existing inter-expert sparsity of the routed experts without significant accuracy loss. We explore intra-expert activation sparsity across six open source MoE models ranging from 1B to 397B parameters, and extend the MoE execution pipeline of vLLM to leverage intra-expert activation sparsity while preserving the existing MoE layer optimizations, achieving up to 2.0x speedup in MoE layer execution and 1.2x end-to-end speedup compared to the dense vLLM baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.