Square-Root Gradient Routing: Geometrical Understanding of Router Optimization in Sparse MoEs
Abstract
Sparse Mixture-of-Experts (MoE) is widely adopted to scale modern large language models. The router activates a small fraction of experts for each token, playing a critical role towards efficient and effective scaling. However, the router's dynamics during MoE training remains poorly understood. In this work, we rigorously study the routing dynamics through a tractable piecewise regression problem, where we simplify the routing with a boundary partition model. Our theory analytically proves that the loss objective is locally cubic in boundary displacement, yielding a quadratically vanishing gradient. Both gradient descent and Adam consequently perform polynomial convergence , while a signed square-root gradient update (SqrtGD) achieves geometric convergence with exponential loss decay . Inspired by the mechanistic understanding, we propose the square-root gradient routing (SR) on MoE-based LLM and evaluate in LLM pretraining. Our experiments on FineWeb demonstrate that the benefits from SR increase with the sparsity of routing adopted in MoE, while its effectiveness is also coupled with the decision on gating mechanism and base optimizer choices. Our findings establish a rigorous mechanistic understanding of the routing dynamics in mixture-of-experts models, which further inspires practical optimizer design for MoE-based foundation models with high sparsity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.