acceptodds
Under review as a conference paper at ICLR 2027

BEYOND OVERLAP: OS-AWARE, BANDWIDTH-CENTRIC EXPERT STREAMING FOR MOE INFERENCE ON UNIFIED-MEMORY SYSTEMS

Abstract

Mixture-of-Experts (MoE) models reduce per-token computation but can exceed the memory capacity of unified-memory devices, forcing router-selected experts to be streamed from SSD. We show that the resulting bottleneck is not expert computation or unified-memory bandwidth, but the realized bandwidth of the OS-mediated SSD-to-unified-memory path. XSTREAM increases this bandwidth by issuing routed-expert requests concurrently, jointly tuning request granularity and I/O/decompression parallelism, and streaming each expert through lossless decompression, the Gate and Up projections, and the dependent Down projection without waiting for the whole expert. On an NVIDIA GB10 system, it achieves 2.07–4.65× higher throughput than ZipMoE at matched memory budgets, and 1.47–1.53× higher throughput than serial one-expert-ahead execution without pipelining. These results demonstrate that, when experts must be fetched, exposing and consuming storage parallelism is a first-order requirement for SSD-backed MoE inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.