Shared Router Dynamics in Sparse MoEs: From Canonical Geometry to Expert Caching
Abstract
Sparse mixture-of-experts (MoE) models use an independently parameterized router at each sparse layer, yet their routing states exhibit substantial predictability across depth. We ask whether this predictability reflects reusable structure that is obscured by layer-specific coordinate systems. We isolate each router's control subspace and show that its factorization has an exact orthogonal gauge symmetry, which defines a principled quotient space for cross-layer comparison. We then fix this gauge globally with generalized orthogonal Procrustes analysis before fitting any transition model. Across four MoE architectures, a single shared transition recovers 79–90% of the held-out of independently fitted layer-specific dynamics, and the canonical forecasting advantage replicates on a second corpus across multiple architectures. Although generic residual representations are also predictable across depth, matched-rank comparisons show that canonical router coordinates preserve substantially more information about expert selection. Direct interventions nevertheless reveal a predictive–causal gap: accurate routing-state predictions need not preserve language-model loss when substituted for native routing. As a downstream application, we use canonical cross-layer forecasts to anticipate later expert demand while leaving native routing unchanged. In physical expert-offloading experiments, forecast-guided caching reduces expert transfers by up to 17.0% and improves token throughput by up to 5.7%, while preserving native routes and outputs. These results provide evidence for reusable routing geometry and dynamics across depth and show that the same structure can support MoE memory management without replacing the router itself.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.