Concentration-Scaled Grassmannian Routing for Controllable Mixture-of-Experts
Abstract
A sparse mixture-of-experts model is useful only if its router can specialize experts without starving them and can adapt inference cost after training. Standard top- routers expose the first problem during training but provide only an indirect post-hoc compute interface. We propose Grassmannian MoE (GrMoE), which represents each expert by an orthonormal subspace and scores a normalized token by concentration-scaled projection energy. The resulting affinity is basis invariant and bounded, while a global concentration scale controls the sharpness of the routing distribution from one checkpoint. We distinguish these geometry-specific properties from generic consequences of softmax scaling and give a conditional load-balance bound under explicit mixture-balance and effective-margin assumptions. At 350M parameters, GrMoE improves perplexity and severe-concentration frequency over standard top- routing; after matching descriptor parameters, scoring MACs, diversity regularization, expert compute, and realized fan-out, it retains a smaller -PPL residual over the strongest Euclidean control. A fixed checkpoint reaches two active experts across with only PPL spread, and downstream, /, second-family, expert-parallel, and memory studies preserve the same scoped picture. We therefore present GrMoE as a geometric routing primitive for stable one-checkpoint compute control, not as a universal or frontier-scale routing claim.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.