High-Dimensional Dynamics of Mixture-of-Experts: Phase Transitions, Routing, and Specialization
Abstract
Mixture-of-Experts (MoE) architectures have been pivotal in scaling modern language models, yet the coupled evolution of their routing mechanisms and expert parameters remains underexplored. We give a high-dimensional analysis of the training dynamics of an MoE layer with experts in the feature-learning regime, using dynamical mean-field theory (DMFT) to derive a closed system of integro-differential equations for the router and expert order parameters. The routing nonlinearity yields two differences from single-pathway DMFT: a non-Gaussian effective action, which we close by treating the routing fields as Gaussian, and a router-specific Onsager correction, needed because the router is trained on the examples it routes. We validate the equations against gradient flow at , and and show that their error grows as the routing temperature falls. For softmax routing we identify the gate with a Random Energy Model in a signal-dependent field, and use it to separate two distinct processes: a dynamical onset, where a symmetric saddle amplifies the initialization bias, and a thermodynamic condensation of the router. In gradient flow, cold routers condense onto poorly aligned experts while warm routers stay diffuse and align, and the specialization reached grows with the initialization bias, which in our longest runs is amplified by a finite gain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.