acceptodds
Under review as a conference paper at ICLR 2027

ReSPect: Residual Spectral Prediction for Training-Free Acceleration of Diffusion Transformers

Abstract

Diffusion transformers underlie current image and video generation models, but incur substantial sampling costs due to iterative denoising. Training-free caching accelerates sampling through two coupled decisions: which steps to evaluate, and what to substitute for the skipped features. Many existing schedules gate on raw distances between adjacent features, a drift signal with poor signal-to-noise ratio, where stochastic variation masks meaningful drift, while ignoring how the noise level itself evolves along the sampling trajectory. Many existing reconstructors, in turn, rely on local approximation, copying the most recently computed features or extrapolating from a short local history. This approach causes approximation errors to increase with the skip length and sample quality to degrade at high speedups. We propose Residual Spectral Prediction for Efficient Caching of Transformers (ReSPect), a training-free cache that jointly addresses these limitations through spectrally aligned reuse decisions and residual trajectory prediction. Scheduling decisions are based on a spectrally aligned representation obtained from a timestep-dependent filter derived from the linear minimum mean square error (linear-MMSE) denoising formulation, which preserves content-relevant components while suppressing noise. Skipped residuals, rather than hidden states, are forecast by a weighted sum of ridge-regularized Chebyshev regression over the accumulated residual history, which captures the global trajectory, and a first-order discrete Taylor correction, which captures local dynamics. Analysis shows the forecast error admits a skip-length-independent bound, whereas the corresponding bounds for copy-reuse and Taylor extrapolation grow with skip length. Across extensive experiments on FLUX.1-dev, Wan2.1-1.3B, and HunyuanVideo, ReSPect achieves speedups of , , and over full denoiser evaluation while achieving lower output error than existing caching baselines at matched or lower compute.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.