acceptodds
Under review as a conference paper at ICLR 2027

Rethinking MoE Routing: Unlocking the Sparsity Potential of Pretrained Models

Abstract

Mixture-of-experts (MoE) models expand capacity through sparse activation and are widely used in large language models. However, fixed Top- routing assigns every token the same expert budget, creating redundant computation. Training-free skipping based on routing scores or importance estimates can compromise quality under aggressive pruning, while jointly adapting routers and experts requires optimizing a large pretrained parameter set. Our analysis reveals that pretrained MoEs can retain their capabilities under smaller expert budgets by relearning expert selection, while adaptive budgets across tokens and layers further reduce computation without sacrificing quality. Guided by these findings, we develop BEAM (Binary Expert Activation Masking), which combines hidden states with a routing prior to select which candidate experts execute. Under a shared sparsity objective, mask routers distribute a global expert budget across tokens and layers, while layer-wise self-distillation preserves the original model's representations and predictions with pretrained weights frozen. The learned policy exhibits dynamic sparsity, retaining computation for sensitive tokens and layers while pruning redundant activations elsewhere. On Qwen3-30B-A3B, BEAM preserves over 99% of the original model's performance while reducing expert activations by 50%. To deploy dynamic MoEs efficiently, we optimize BEAM kernels for high-performance inference engines. With our kernels integrated into vLLM, BEAM achieves up to throughput and TPOT speedups over the original MoE.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.