acceptodds
Under review as a conference paper at ICLR 2027

Less Can Be More: Learning Pruned Top- in SMoE-based LLMs

Abstract

Sparse Mixture of Experts (SMoEs) typically retain the fixed Top- expert budget chosen during pretraining when deployed on downstream tasks. But is this computation budget always necessary or even optimal? Through an oracle analysis across nine reasoning and non-reasoning SMoE models, we find substantial redundancy in expert activation: smaller budgets can preserve correct predictions, while some inputs that fail under the native Top- become solvable when fewer experts are activated. Overall, the oracle reduces the average expert budget by **40.8%** while yielding a **26.8% relative accuracy improvement**. Motivated by these observations, we introduce , which adapts both which experts to favor and how many experts to execute while keeping pretrained model parameters frozen. LEPO combines lightweight expert-preference adaptation with grouped expert-budget policies and learns from the available context through guidance from the original model, without requiring task labels or oracle budget annotations. Across test-time adaptation, offline budget learning, and permanent expert pruning, LEPO consistently improves the accuracy and efficiency trade-off over fixed and adaptive expert-allocation baselines. In the settings, LEPO uses only **75-87%** of the native routed-expert budget while largely preserving or improving accuracy. Our results suggest that downstream Top- should be treated as an adaptive computation-allocation problem rather than a fixed consequence of pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.