acceptodds
Under review as a conference paper at ICLR 2027

Expert-Aware Router: Coupling Routers and Experts in Mixture-of-Experts

Abstract

Sparse Mixture-of-Experts (MoE) models scale Transformer capacity by activating only a small subset of experts for each token. However, standard MoE routers are parameterized independently from the experts they select, so routing scores do not directly reflect evolving expert representations. Existing expert-aware methods strengthen this coupling using expert-side activations, requiring activating many or all experts before routing, introducing substantial training overhead. We propose Expert-Aware Router (EAR), a lightweight routing mechanism that derives each expert's routing representation directly from its gate projection matrix through a shared learnable probe vector. The resulting expert-derived routing matrix is combined with the standard learned router, allowing updates to expert parameters to directly affect routing decisions while retaining flexibility in token assignment. This coupling enables the router and experts to co-adapt during training while preserving the standard linear routing formulation. Across MoE models from 1B to 16B parameters and training budgets up to 500B tokens, EAR consistently improves downstream performance, increasing average accuracy by up to 0.76 points over standard MoE. % Under Long Horizon 16B training, For the 16B model trained for 500B tokens, it reduces the maximum expert-load violation by 71.7%, increases router–expert top- overlap from 19.45% to 21.72%, and substantially reduces inactive neurons. EAR also improves token efficiency, achieving a 1.12 gain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.