acceptodds
Under review as a conference paper at ICLR 2027

MMoE: Multi-Level Routing Capacity Allocation for Extremely Sparse Multimodal MoE Inference

Abstract

Mixture-of-Experts (MoE) has become a prevalent architecture for scaling multimodal large language models (MLLMs), yet activating multiple experts for hundreds of visual tokens can still incur substantial inference overhead. In this work, we investigate routing redundancy in MoE-based MLLMs and identify two complementary forms of heterogeneity: (i) different MoE layers exhibit markedly different sensitivities to routing-capacity reduction, and (ii) visual tokens within the same layer require substantially different amounts of expert computation. Motivated by these observations, we propose MMoE, a multimodal-specific multi-level routing-capacity allocation framework for efficient MoE inference under a fixed visual-token expert budget. MMoE first profiles layer sensitivity through output-distribution perturbation and allocates routing capacity across MoE layers accordingly. It then redistributes each layer's budget among visual tokens by prioritizing token-expert pairs with larger routing weights, thereby preserving more routing mass under the same budget. Experiments on two representative MoE-based MLLMs across eight vision-language benchmarks show that MMoE achieves the highest normalized average performance compared with existing routing strategies across multiple capacity budgets. Under aggressive sparsity, it improves normalized average performance by up to 5.57% over the strongest baseline while skipping more than 91% of visual-token expert activations, and achieves up to 2.62× prefill speedup on Kimi-VL-A3B-Instruct.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.