acceptodds
Under review as a conference paper at ICLR 2027

Budget-Aware Expert Loading for Memory-Constrained MoE Inference

Abstract

Mixture-of-experts (MoE) models only activate few experts per token, yet CPU-offloaded inference can still incur substantial overhead from repeatedly loading missing experts to GPU when its memory is limited. In this paper, we propose , a training-free controller that explicitly coordinates expert loading under a fixed GPU cache capacity and a soft model-wide loading budget. The controller values each load by its marginal router-mass recovery over resident alternatives and compares this gain against a shared price updated by a request-local Lyapunov queue. Coupled with global LRU eviction, the controller coordinates loading across layers and decoding steps rather than loading every missing native-route expert. From a theoretical perspective, we prove an upper bound on the total number of loadings per request, thereby ensuring that our controller is time-efficient. We then conduct extensive experiments, involving three MoE models and four tasks spanning reasoning and code generation. The results demonstrate that substantially reduces expert-weight traffic and improves decode throughput, while achieving a comparable performance. For example, on MATH L5, using Lyapunov-LRU to run Qwen3-30B-A3B on L20 achieves a speedup of 1.38× and only incurs a tiny 4.4 percentage-point accuracy drop.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.