Optimal Hyperparameter Scaling Laws Across MoE Sparsity
Abstract
Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes hyperparameter transfer challenging. In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: optimal learning rate and batch size vary with activation ratio, and neither total nor activated parameter count explains these shifts. To characterize this dependence, we conduct 1,800 pre-training runs spanning six activated-parameter scales and models with up to 6B total non-embedding parameters, processing approximately 20 trillion tokens at a cost of 200,000 equivalent H800 GPU-hours. Our results reconcile conflicting findings in prior work by revealing two scaling regimes. At fixed sparsity, optimal batch size follows a power law in training tokens , whereas optimal learning rate scales with training compute and remains robust to the allocation between model size and data. Across sparsity levels, activation ratio enters both relationships as a multiplicative power-law factor, yielding unified laws that transfer across MoE sparsity levels. Large-scale evaluation shows that the scaling form outperforms alternative functional forms. On a held-out ultra-sparse MoE with 12B total parameters and only 1/64 of its experts activated, the predicted hyperparameters remain close to the observed optima, supporting joint extrapolation across model scale and sparsity. Further experiments demonstrate transfer across expert granularities and isolate the effect of activation ratio from total expert count.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.