acceptodds
Under review as a conference paper at ICLR 2027

SEP: Speculative Expert Prefetching for Low-Latency Mixture-of-Experts Inference

Abstract

Mixture-of-experts (MoE) decode is memory-bandwidth bound: each token generation step must stream expert weight matrices from HBM before the expert GEMM can execute. We introduce Speculative Expert Prefetching (SEP), which predicts each token's active experts from the pre-attention hidden state and prefetches expert weights overlapped with attention, routing, and multi-GPU communication. SEP rests on a hardware-algorithm co-design insight, that on-package cache holds more experts than the router's per-GPU top-, enabling overprovisioned prefetching that relaxes the prediction objective from exact expert prediction to per-GPU coverage. When this overprovisioning is large, spare cache slots tolerate mispredictions and a lightweight predictor suffices, while tight cache capacity demands higher accuracy. A second limit is temporal, as the predictor and the expert weight transfer must together complete before the expert GEMM begins. We evaluate 18 predictor architectures across four modern MoE models, demonstrating strong out-of-distribution generalization across diverse task categories. Our best predictor achieves - accuracy in predicting experts, translating to - decode-GEMM speedups at batch 1 on production MoE kernels.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.