acceptodds
Under review as a conference paper at ICLR 2027

TEMPORALLY CONSISTENT ROUTING: A TRAINING- TIME OBJECTIVE FOR CACHE-EFFICIENT MIXTURE-OF- EXPERTS

Abstract

Mixture-of-Experts (MoE) models activate only a sparse subset of experts per token, yet consecutive tokens frequently activate different experts, causing constant weight swapping between slow storage and fast memory on edge devices. Existing remedies are either system-level (such as caching heuristics) or post-hoc (such as router fine-tuning), leaving the root cause unchanged during pretraining. We propose a differentiable routing consistency loss that penalizes abrupt expert switches between adjacent tokens, encouraging the router to maintain the same expert assignment across semantically coherent spans. The method requires no architectural changes and adds only a single hyperparameter, λ. Unlike post-hoc approaches, it allows expert representations and routing decisions to co-adapt from the first training step. Experiments on small and medium-sized MoE language models show that the method reduces the expert switch rate by up to 59% while simultaneously improving perplexity on the medium model. It also reduces cache misses by up to 3.92×, Pareto-dominating post-hoc fine-tuning on the quality-locality frontier. These results suggest that routing temporal locality is most effectively learned during training rather than added afterward.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.