Beyond Forward Counts: Caching in MoE Diffusion Models
Abstract
Caching reduces repeated token computation in diffusion language models, but its serving benefit in offloaded mixture-of-experts (MoE) models depends on how reuse propagates through expert execution and weight movement. These operations occur at different granularities: routed query positions are grouped into expert calls, while expert residency determines which weights must move. Consequently, a large reduction in processed query positions need not produce a comparable reduction in expert calls or transferred bytes, and forward count alone cannot characterize serving efficiency. We introduce TRACE-MoE, an execution-aware evaluation framework that jointly measures query visits, forward evaluations, expert calls, transferred bytes, and request latency to connect token reuse with its downstream execution costs. Family-matched controls and paired timing compare cache policies under identical requests and runtime configurations, while independent diagnostic runs reconcile the measured work counters. Workload-preserving controls provide a timing reference, and fixed-forward-count comparisons reveal changes in work within denoising steps. Controlled executions of Fast-dLLM and dKV-Cache on LLaDA2 mini checkpoints compare cache policies with their corresponding family controls at a common 64-token generation budget. These measured configurations include lower latency with more forward evaluations, with paired savings of up to 59.0% when forward count increases. TRACE-MoE makes the relationships among token reuse, expert demand, weight movement, and latency explicit, providing a practical basis for evaluating cache policies under matched execution conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.