Expert-Guided Router Recovery for Stable Mixture-of-Experts Pre-training
Abstract
Mixture-of-experts (MoE) language models expand model capacity while limiting per-token computation through sparse expert activation. However, loss spikes during from-scratch MoE pre-training remain a critical yet underexplored failure mode. To address this, we propose Expert-Guided Router Recovery (EGRR), a pre-training recovery method that bridges the gap between router decision-making and expert execution by enabling the router to assign tokens to experts according to capability information derived from the experts themselves. EGRR is based on the insight that around loss spikes, router-selected experts exhibit unusually low activation magnitudes, suggesting limited capability for processing the current representations. In EGRR, once a loss spike is detected, EGRR rolls training back to the latest stable checkpoint and densely activates all experts to guide router tuning with expert-derived capability information, while all non-router parameters are frozen. After 100M tokens recovery period, dense expert activation is removed and standard MoE pre-training resumes. EGRR therefore introduces additional computation only during recovery, with no additional inference cost. We evaluate EGRR in a 100B tokens from-scratch pre-training run of a 700M parameters MoE language model, demonstrating improved training stability and downstream performance over standard MoE training. The code is available at https://github.com/saz97/Expert-Guided-Router-Recovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.