VeloCache: Frequency-Decoupled Velocity Caching for Training-Free Video Diffusion Acceleration
Abstract
Training-free step caches accelerate video diffusion by reusing cached computation at selected denoising steps. Prior work mostly refines when to reuse; what is reused is typically the whole cached signal, which can wash out texture, or per-block feature forecasts whose cost grows with model size. Analyzing the per-step spectrum of the denoiser output, we find a volatility crossover: about one fifth into the trajectory, changes in the low spatial band that carries coarse structure decline, while the high band that carries fine texture keeps changing. This motivates two decisions. What to reuse: VeloCache retains the cached low band and advances the high band by a first-order trend from the two most recent computed outputs. Only true evaluations refresh this history, so skipped steps never feed predictions back into the cache. A local error analysis gives the condition under which the correction improves on replay, and shows how unequal sampling intervals limit it. When to evaluate: an early segment set by the pooled crossover is computed in full and the remaining budget is spread uniformly, yielding one shared rule across backbones. VeloCache reads no internal features and stores two outputs per guidance stream. At comparable evaluation budgets, it reaches the lowest LPIPS to the full-step reference on Wan2.1 and CogVideoX ( vs. on Wan2.1), competitive LPIPS on Cosmos3-Nano, and the highest VBench quality of the compared caches on all three backbones. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.