acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Timestep-level Cache in Diffusion: A Minimalist Framework for Effective Step Scheduling and Error Rectification

Abstract

Timestep-level caching has emerged as a prominent technique for accelerating diffusion transformers due to its simplicity and strong empirical performance. Although existing approaches are predominantly formulated as residual caching (i.e., model output minus input), we show through first-order Taylor expansion that these formulations are equivalent to direct output caching and can therefore be interpreted as enlarging the numerical integration step size. This interpretation recasts timestep-level caching as a numerical integration problem, where generation quality is determined by both when larger integration steps are taken and how the resulting truncation errors are controlled. Consequently, we identify two fundamental factors governing the effectiveness of timestep-level caching: the allocation of cached steps and the rectification of cache-induced errors. Motivated by these observations, we propose a minimalist yet highly effective acceleration framework that jointly addresses these two challenges. Specifically, we introduce a quality-driven scheduler to determine cache step allocation and a lightweight neural compensator to correct truncation errors along the denoising trajectory. Unlike existing approaches that rely on heuristic schedules and static linear corrections, our framework allocates cache steps according to their impact on final image quality and employs a neural compensator that better adapts to complex denoising trajectories, thereby achieving superior quality-acceleration trade-offs, all with minimal training overhead. Extensive experiments on FLUX.1-dev demonstrate that our method substantially outperforms existing baselines, achieving an LPIPS of 0.1914 at 10241024 resolution while delivering a 2.77 inference speedup. As an architecture-agnostic framework, it further generalizes to video diffusion models, achieving an LPIPS of 0.0779 with a 2.91 speedup on CogVideoX-2b. Notably, across all evaluated models, the entire pre-processing stage requires less than 30 minutes, highlighting the efficiency and practicality of our approach.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.