Learning Token-Population Geometry for Sparse Mixture-of-Experts
Abstract
Routing in sparse mixture-of-experts (MoE) models determines both token-level computation and the populations from which experts learn. Load balancing controls how much traffic each expert receives, but does not specify how these populations are organized in representation space. We introduce the Geometry-Guided Learnable Router (GGLR), which uses expert input means maintained by an exponential moving average as an explicit population reference for a task-trained router. This reference defines population-relative coordinates for expert selection and target directions whose changes guide router updates, while selection-only biases regulate utilization. We analyze how these population-relative coordinates transform pairwise routing logits and how historical averaging affects the accuracy of the population reference. In matched 1.7-billion-parameter experiments on Nemotron-CC-v2, GGLR achieves the lowest observed endpoint loss and global load imbalance, while producing more separated directions of expert input means and more distinct learned router directions. Compared with Loss-Free Balancing with normalized output gains, GGLR reduces global maximum overload by 23.1%, while also lowering validation perplexity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.