DyS-MoE: Virtualizing Dynamic Expert Execution with Static Windows for Resource-Constrained MoE Inference
Abstract
Mixture-of-Experts (MoE) inference under limited GPU memory relies on expert offloading. Dynamic execution exploits heterogeneous resources but incurs frequent transfers and scheduling overhead, while static expert placement simplifies execution at the cost of GPU utilization. We present DyS-MoE, which uses selective expert replacement to serve dynamic demand through reusable execution windows. Its core abstraction, a Virtualized Static Execution Graph, preserves operations and dependencies while expert bindings change. Slotted Grouped Execution provides stable GPU cache and execution slots, and Importance-Aware Miss Handling maps residency misses and capacity overflow onto these resources while reserving CPU exact computation for sensitive contributions. Together, these mechanisms enable CUDA Graph replay under dynamic expert offloading. On an RTX 4090 at batch 8, DyS-MoE achieves 431.99 and 551.06 decode tokens/s on Qwen3-30B-A3B and DeepSeek-V2-Lite-Chat, respectively, reaching 2.28× and 6.90× the throughput of state-of-the-art baselines while maintaining competitive task accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.