acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Expert Prefetching for Offloaded Diffusion MoE Inference

Abstract

Diffusion mixture-of-experts (MoE) language models repeatedly process the same positions across denoising steps, making expert routing highly stable over time. This makes short-horizon expert prefetching a natural way to hide host-to-device transfer latency. We find that the same stability also limits what prefetching adds beyond a recency-based cache: when the GPU cache can retain most of one denoising step's expert working set, many experts predicted for the next step are already resident under LRU. Motivated by this observation, we separate predictive cache placement from transfer-latency hiding and use a demand-driven asynchronous offloading design: expert residency follows observed demand through a per-layer LRU cache, while cache misses are transferred asynchronously and overlapped with computation, with no speculative prefetching. This isolates the additional value of prediction beyond cache recency and asynchronous demand loading. Deterministic miss accounting shows that, near and above the working-set scale, last-step prefetching reduces few or no demand misses beyond LRU while still introducing speculative transfers; in contrast, an oracle with true future routing cuts misses by . Bandwidth throttling further shows that these speculative transfers compete with demand traffic for host-to-device bandwidth. Across LLaDA-2.0-mini and the larger LLaDA-2.0-flash on an A5000 and an H200, our demand-driven asynchronous design improves throughput by – over asynchronous last-step prefetching near and above the per-step working-set scale, with all paired 95% confidence intervals excluding zero. The same ordering holds across five tasks and batch sizes 1–8. The effect depends on cache capacity: the no-prefetch advantage is significant for –, but falls to (95% CI ) at , where the cache can no longer retain the recent working set. An interval-refresh placement policy inspired by TIDE does not reduce misses below LRU in our execution setting. Head-to-head against TIDE's released implementation at the same expert budget on the same GPU, our system achieves a speedup ( at half the budget). These results show that expert prefetching in offloaded diffusion MoEs should be evaluated by the additional misses it avoids beyond cache residency and asynchronous miss handling, not by routing predictability alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.