acceptodds
Under review as a conference paper at ICLR 2027

MoRE: Scaling mixture of experts with hardware-aware low-rank routing

Abstract

Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with experts and hidden dimension , its per-token cost dominates the MoE layer once is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank and reduces the routing cost to . We prove that rank logarithmic in suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q&A benchmarks after pretraining, while matching reasoning ability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.