StaVER: A Probabilistic Formulation for Stable Routing in Mixture-of-Experts
Abstract
Mixture-of-Experts (MoE) models scale capacity efficiently by routing each token to a small subset of experts. A key challenge in MoE models is routing instability, which can reduce training efficiency and impair model performance. Through empirical studies and theoretical analysis, we identify the discontinuity of hard Top- selection as a source of this instability: perturbations can abruptly change the selected expert set, leading to performance degradation. To address this, we propose StaVER, a unified Bayesian framework that formulates MoE routing as a probabilistic inference problem. StaVER models expert assignment as a latent variable and performs routing by sampling experts from its distribution. Under this probabilistic formulation, the expected expert-selection frequencies vary continuously to perturbations, smoothing the abrupt changes in expert selection induced by hard Top- routing. To compute this intractable posterior distribution over expert assignment, we follow variational inference to derive a unified training objective. This yields a prior-matching term that naturally promotes load balancing, thus eliminating the need for an additional heuristic auxiliary loss. Through extensive experiments on MoE models ranging from 1.8B to 6.6B parameters, we demonstrate that StaVER achieves more stable routing, more balanced expert utilization, and consistent performance gains on both upstream and downstream tasks. Across eight downstream benchmarks, StaVER improves average accuracy up to 2.6 percentage points over the strongest baseline. In the pre-training evaluation, it obtains the lowest PPL, while reducing estimated training FLOPs by approximately 18%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.