Fine-Grained Caching for Diffusion Transformers with Few Calibration Conditions
Abstract
Diffusion transformers require repeated denoiser evaluations, making image and video generation computationally expensive. We propose a training-free framework that uses a few calibration conditions to construct a fixed, fine-grained module-reuse schedule without schedule search. The design is motivated by an empirical effect along deterministic sampling trajectories: with initial noise fixed, condition-dependent deviations in several module outputs evolve similarly across adjacent steps, so temporal differencing attenuates much of their variation. We term this effect Conditional Common-Mode Rejection (CCMR). Motivated by this observation, we rank timestep–layer–module locations using normalized ratios of consecutive feature displacements and construct a Cache Book through two-stage calibration. The second stage updates a scorer-only reference with each provisional bit after recording its indexed score, while retaining full-compute trajectories. At inference, the fixed schedule requires no per-input policy estimation and can be nested within fixed step-wise cache policies. Module caching alone achieves – speedup on the evaluated image models and – on the video models. Combining it with MagCache or SeaCache provides further acceleration at a model-dependent cost in paired fidelity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.