acceptodds
Under review as a conference paper at ICLR 2027

Know Before You Route: Dual-Signal Expert Prefetching with Conformal Coverage Guarantees for Sparse MoE LLMs

Abstract

Sparse Mixture-of-Experts (MoE) LLMs achieve high capacity with sparse activation, but under GPU memory constraints the full expert pool may not fit on device, forcing expert-weight transfers from host memory onto the decode critical path. We introduce CEPA, a retraining-free dual-signal expert prefetching framework. An early velocity-corrected cross-layer proxy initiates one likely-expert transfer, while a later signal evaluated on the non-extrapolated layer-boundary residual constructs a calibrated expert set. The boundary residual remains a proxy for the post-attention router input, yet corrected lower-tail conformal calibration provides layerwise marginal whole-active-set coverage when calibration and test score–set pairs are exchangeable. CEPA speculates only on transfer timing: exact routing and expert computation are preserved by fetching every missing active expert on demand. Across four sparse MoE LLMs spanning 14.3B–35B total parameters under a 12 GB GPU memory budget, CEPA reduces decode latency by up to and retains a speedup against a matched CUDA-Graph baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.