DART: Diffusion Acceleration via Rectified Temporal Quantization
Abstract
Diffusion models achieve high generative quality but incur substantial computational cost from iterative denoising. Feature caching and low-bit quantization each reduce this cost, and combining them promises multiplicative speedups. Yet we find that directly stacking the two strategies is unstable: cache reuse introduces a temporal state drift along the denoising trajectory, while quantization distorts the denoising operator, and the two errors interact in a way that local feature distance fails to predict. We propose DART, a training-free joint cache–quantization framework that addresses two questions explicitly—when to recompute and how to use cached quantized states—under a unified trajectory-sensitivity view. DART first measures a relative update sensitivity proxy along the sampling trajectory and partitions the timesteps into segments whose sensitivity fluctuation is bounded; the partitioning, formulated as a constrained dynamic program, decides where caching is safe. Within each segment, DART rectifies the cached quantized states through dual-domain affine correction, which decouples cache-induced input shift from quantization-induced operator distortion, and shares a segment-wise quantizer that adapts to temporal heterogeneity. The rectification parameters are absorbed into adjacent linear and dequantization operators whenever the required affine conditions hold, yielding negligible deployment overhead. Across class-conditional generation on ImageNet-256 with LDM and DiT-XL/2 and text-to-video generation on HunyuanVideo, DART consistently improves the quality–efficiency frontier over caching-only, quantization-only, and joint cache–quantization baselines under matched-speed and matched bit-width settings. For instance, on LDM-4 with W8A8 quantization DART improves FID from 4.03 to 3.83 at 8.12x versus CacheQuant's 7.87x, and improves FID from 7.21 to 6.53 at the most aggressive 18.93x setting where directly stacking DeepCache and W8A8 collapses to FID 15.36, an 8.83 FID gap at the same speed. On DiT-XL/2 at 50 sampling steps DART further improves FID from 5.43 to 5.08 at 13.2x versus the 12.7x of the recent joint baseline Q&C, and on HunyuanVideo it lowers LPIPS by 23 from 0.189 to 0.145 and lifts VBench from 80.35 to 80.42 at 4.12x versus DPCache's 4.05x.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.