acceptodds
Under review as a conference paper at ICLR 2027

Fine-Grained Caching for Diffusion Transformers with Few Calibration Conditions

Abstract

Diffusion transformers require repeated denoiser evaluations, making image and video generation computationally expensive. We propose a training-free framework that uses a few calibration conditions to construct a fixed, fine-grained module-reuse schedule without schedule search. The design is motivated by an empirical effect along deterministic sampling trajectories: with initial noise fixed, condition-dependent deviations in several module outputs evolve similarly across adjacent steps, so temporal differencing attenuates much of their variation. We term this effect Conditional Common-Mode Rejection (CCMR). Motivated by this observation, we rank timestep–layer–module locations using normalized ratios of consecutive feature displacements and construct a Cache Book through two-stage calibration. The second stage updates a scorer-only reference with each provisional bit after recording its indexed score, while retaining full-compute trajectories. At inference, the fixed schedule requires no per-input policy estimation and can be nested within fixed step-wise cache policies. Module caching alone achieves – speedup on the evaluated image models and – on the video models. Combining it with MagCache or SeaCache provides further acceleration at a model-dependent cost in paired fidelity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.