acceptodds
Under review as a conference paper at ICLR 2027

Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Execution

Abstract

Mixture of Experts (MoE) architecture has become the standard for state-of-the-art large language models, owing to its computational efficiency through sparse expert activation. However, sparsity through finer expert granularity is becoming increasingly difficult to achieve due to fundamental training challenges such as expert collapse and load imbalance. In this work, we explore and leverage intra-expert activation sparsity as a complementary and underexplored dimension of sparsity in MoE models. Surprisingly, substantial intra-expert sparsity is readily available in off-the-shelf pre-trained MoE models, providing up to 98% sparsity on top of existing inter-expert sparsity of the routed experts without significant accuracy loss. We explore intra-expert activation sparsity across six open source MoE models ranging from 1B to 397B parameters, and extend the MoE execution pipeline of vLLM to leverage intra-expert activation sparsity while preserving the existing MoE layer optimizations, achieving up to 2.0x speedup in MoE layer execution and 1.2x end-to-end speedup compared to the dense vLLM baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.