RoDCache: Routing-Drift-Aware Caching for Accelerating Sparse MoE Diffusion Models
Abstract
Diffusion models incur substantial inference cost due to repeated execution of the denoising network across multiple timesteps. Caching methods reduce this cost by reusing results from previous denoising steps based on estimates of cache validity. However, in sparse MoE diffusion models, existing cache-validity signals overlook how routing drift in expert assignments and routing weights alters the active computation path and affects reuse reliability. To address this limitation, we present RoDCache, a training-free caching method that accounts for routing drift in cache-validity assessment. It comprises two components: Routing Drift Quantification (RDQ) and Routing-Drift-Constrained Scheduling (RDCS). RDQ quantifies model-level routing drift by aggregating changes in expert assignments and routing weights across tokens and MoE layers. Based on this signal, RDCS selects the longest reuse interval whose predicted routing drift does not exceed a tunable threshold. Across three sparse MoE diffusion models, \method achieves – speedups and consistently outperforms existing caching methods in PSNR, SSIM, and LPIPS, with PSNR gains of up to dB. Further analysis shows that RoDCache adapts computation allocation to prompt difficulty and improves quality consistency across prompts. Code will be released at https://anonymous.4open.science/r/RoDCache.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.