acceptodds
Under review as a conference paper at ICLR 2027

One Threshold Is Not Enough: Prompt-Invariant Caching Schedules for Video Diffusion

Abstract

Video diffusion models remain expensive because denoising repeatedly evaluates a costly model backbone. Dynamic caching methods like TeaCache, EasyCache, and DiCache reduce this cost by reusing prior outputs while an accumulated drift signal remains below a threshold . Although existing methods improve the drift signal, they hold fixed throughout denoising. We show that fixed thresholds are structurally suboptimal: varying across denoising achieves quality-latency tradeoffs unattainable by any fixed threshold across three caching methods with distinct drift signals. We further show that, surprisingly, each method's drift trajectory is nearly prompt-invariant across seven method-model combinations. The shape of the drift trajectory is determined by the model and caching method; therefore, a threshold schedule can be calibrated offline and reused across generations. Based on these findings, we present ACID (Adaptive Caching for vIDeo generation), which uses a low threshold during critical regions where drift changes rapidly and a high threshold across stable regions. ACID identifies these regions offline from the drift signal's second derivative, requires no training, and adds negligible runtime overhead. Across TeaCache, EasyCache, and DiCache on HunyuanVideo, Wan 2.1, and CogVideoX, ACID pushes the speed-quality Pareto frontier beyond fixed thresholds. On TeaCache with HunyuanVideo, it achieves speedup over no caching and 38% additional speedup over a conservative fixed threshold, with less than 0.3 dB PSNR, 0.01 SSIM, and 0.01 LPIPS degradation relative to that fixed-threshold configuration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.