acceptodds
Under review as a conference paper at ICLR 2027

Temporal Mixture-of-Experts: Serving Memory that Scales with Active Parameters

Abstract

Sparse Mixture-of-Experts (MoE) models activate only a small fraction of their parameters per token, yet serving one still requires the whole expert pool in fast memory, making inference impractical on memory-constrained local devices. We introduce Temporal MoE, a new routing policy that reuses all but one of the previous token's experts for each layer. This allows us to keep only the active experts in RAM instead of the entire pool, and schedule loading the incoming expert from flash to overlap with computing the reused ones, effectively hiding transfer overhead. When pre-trained from scratch, Temporal MoE reduces serving memory by while preserving 72–82% of the standard MoE advantage over dense models in cross-entropy and downstream accuracy. This scaling trend holds consistently across compute-optimal budgets from to FLOPs. Probing routers demonstrate that the temporal routers focus more on context rather than token identity and spread each expert's use more evenly across the token stream. In real-device hardware benchmarks across an RTX A6000, an Intel laptop, and a Pixel 10a where the model otherwise does not fit, an 11B-scale model keeps 66–83% of the standard MoE's decode speed, and the workstation realizes the reduction. Applying this constraint post-hoc to released large instruct MoEs demonstrates remarkable robustness (holding especially well for sparser models). Moreover, just two GPU-hours of distillation recovers most of the accuracy lost under the constraint for Gemma4-26B and Qwen3.5-35B. These results show that Temporal MoE can move the quality–memory frontier for local inference by making expert reuse a core part of both training and serving.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.