acceptodds
Under review as a conference paper at ICLR 2027

Control-MoE: Mixture of Experts with Controllable Dynamic Sparsity

Abstract

Mixture-of-Experts (MoE) has become a foundational approach for scaling large models by enabling selective activation of parameters within the multi-layer perceptron (MLP) layers. However, traditional expert selection employs a single router per layer, often leading to inflexible activation rate. Furthermore, the use of a predetermined number of experts per token constrains the adaptability of network sparsity. We propose One Choice at A Time (OCAT), an innovative routing strategy that utilizes parallel Top-1 routers for expert assignment in each layer. By integrating a skip mechanism into each router, which allows bypassing when predictive precision is maintained, OCAT determines both the relevant experts and their ideal count. To automatically regulate the overall sparsity, we further incorporate a Model Predictive Control (MPC) framework that dynamically adjusts the skip-loss weight to steer the realized skip probability toward a target setpoint without manual tuning. Comprehensive evaluations demonstrate that OCAT substantially improves both model performance and operational efficiency. Comprehensive evaluations demonstrate that OCAT substantially improves both model performance and operational efficiency, with OCAT models achieving an average 2% improvement over Top-K router baselines at comparable activation levels across different model scales. It further exhibits test-time scaling capability, allowing the average number of activated parameters to be dynamically adjusted.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.