acceptodds
Under review as a conference paper at ICLR 2027

Sparser-MoE: Beyond Expert Sparsity with Heterogeneous Neuron-Level Sparsity

Abstract

Large language models (LLMs) have demonstrated strong performance across diverse domains, but their high inference costs remain a challenge. To address this issue, prior work has introduced methods that selectively activate neurons in Gated-MLP blocks based on intermediate activation magnitudes, reducing computational costs while largely preserving model performance. However, these methods do not account for the distinctive characteristics of mixture-of-experts (MoE) models, limiting their ability to preserve performance when applying activation sparsification to MoE. We propose Sparser-MoE, an MoE-tailored neuron sparsification framework featuring dynamic expert-wise sparsity and contribution-scale-aware neuron scoring. Motivated by the observation that higher-ranked experts contain more important neurons, Sparser-MoE performs global neuron selection across the top- experts, allowing heterogeneous sparsity to emerge naturally across experts. It further incorporates routing weights into neuron scores to better preserve the scale of the MoE layer output. Together, these designs identify a more effective set of active neurons than existing methods and thereby better preserve the original model performance under sparsification. Experiments across diverse models, model scales, and benchmarks show that Sparser-MoE retains more of the dense model's original performance than existing sparsification methods on nearly every individual benchmark at each evaluated sparsity level, while achieving the highest average benchmark performance at every evaluated sparsity level. To our knowledge, Sparser-MoE is the first Gated-MLP sparsification approach to jointly introduce MoE-specific neuron scoring and selection strategies, and we believe it can serve as a starting point for related research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.