Trust the Router: Adaptive Rank-Space Sparsification for Efficient Mixture-of-Experts
Abstract
Mixture-of-Experts (MoE) models already learn to rank experts through their pretrained routers, yet existing expert-reduction methods often underutilize this existing routing knowledge when reducing computation. We argue that efficient MoE inference should leverage the expert-ranking structure encoded by the pretrained router, rather than selecting experts without explicitly preserving the ranked prefix structure. The central question is no longer merely which experts to select, but how many of the top-ranked experts should actually be executed for each token. Based on this insight, we propose TRUST-MoE, a coarse-to-fine framework that preserves the ranking structure provided by the pretrained router while learning adaptive expert execution. In the first stage, we introduce continuous rank-wise masks together with multi-level knowledge distillation, gradually adapting the original Top- model toward a shorter Top- routing prefix instead of abruptly truncating lower-ranked experts. In the second stage, with the adapted model frozen, a lightweight boundary predictor estimates the counterfactual utility of the next-ranked expert and dynamically decides whether it should be restored for each token-layer pair. This reduces dynamic expert allocation to a simple truncate-or-restore decision, without searching over arbitrary expert subsets. In this way, the MoE itself determines whether a token requires Top- or Top- computation. Experiments across multiple MoE model families and downstream benchmarks show that TRUST-MoE generally improves performance over static expert-reduction and dynamic expert-skipping baselines while maintaining improved inference efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.