SEP: Speculative Expert Prefetching for Low-Latency Mixture-of-Experts Inference
Abstract
Mixture-of-experts (MoE) decode is memory-bandwidth bound: each token generation step must stream expert weight matrices from HBM before the expert GEMM can execute. We introduce Speculative Expert Prefetching (SEP), which predicts each token's active experts from the pre-attention hidden state and prefetches expert weights overlapped with attention, routing, and multi-GPU communication. SEP rests on a hardware-algorithm co-design insight, that on-package cache holds more experts than the router's per-GPU top-, enabling overprovisioned prefetching that relaxes the prediction objective from exact expert prediction to per-GPU coverage. When this overprovisioning is large, spare cache slots tolerate mispredictions and a lightweight predictor suffices, while tight cache capacity demands higher accuracy. A second limit is temporal, as the predictor and the expert weight transfer must together complete before the expert GEMM begins. We evaluate 18 predictor architectures across four modern MoE models, demonstrating strong out-of-distribution generalization across diverse task categories. Our best predictor achieves - accuracy in predicting experts, translating to - decode-GEMM speedups at batch 1 on production MoE kernels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.