acceptodds
Under review as a conference paper at ICLR 2027

A Trainable Memory-Movement Objective for Mixture-of-Experts Decoding

Abstract

Expert-weight transfers can bottleneck batch-one decoding of offloaded mixture-of-experts models. We introduce a differentiable surrogate for expert-weight traffic based on capacity-constrained soft residency, with memoryless and stateful variants. We evaluate learned routing workloads using least-recently-used (LRU) and offline-optimal (Belady) cache misses. On 64-expert models with 3.6–6.8B parameters, training reduces LRU traffic by 14–16× when 25% of expert blocks fit in fast memory, with a 7–13% relative increase in perplexity. A trace-driven H100 replay of the 6.8B model at 12.5% capacity yields 63% higher throughput with blocking transfers and 90% under ideal overlap. Comparisons against a switch-only objective reveal strong dependence on load balancing and cache capacity. At a load-balancing coefficient of 0.01, the memoryless configuration achieves 61% fewer LRU misses at matched mean perplexity on the 414M model with 25% capacity, but not at 12.5%. An analysis of marginal concentration and temporal reuse helps explain these reversals. Increasing the retrofit learning rate substantially narrows the apparent advantage of applying the objective during pretraining, although residual differences remain at matched total compute. These results show how traffic-aware training, deployment budgets, and optimization protocols jointly shape MoE cacheability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.