RACE: Refresh-Anchored Cache for Diffusion Acceleration
Abstract
Flow matching methods have made advanced progress in image and video generation, yet their iterative sampling remains computationally expensive. Cache-based acceleration explores computational redundancy and approximates velocity in certain skipped steps. However, velocity approximation errors accumulate through sampler updates, resulting in the perturbed state for velocity prediction in cache refresh. To improve velocity prediction in cache refresh, the state displacement in the preceding updates should be adjusted. Moreover, we propose that the distribution of state and velocity updates is related along local generation trajectory. This distribution constraint can be easily applied on the estimation of velocity at reuse steps. Building on these insights, we propose **RACE** (Refresh-Anchored Cache), a training-free framework that combines state-guided velocity extrapolation with refresh-anchored integral correction. RACE uses the evolving state to guide velocity prediction and refresh-endpoint errors to compensate for state deviations accumulated during reuse steps. Each refresh step thus serves as an anchor both for subsequent velocity approximation and preceding state correction. RACE is compatible with existing cache schedulers and requires no additional network computation or hyper-parameters tuning. Extensive experiments on FLUX.1-dev, Wan 2.1, HunyuanVideo, and MiniMax-H3 demonstrate that RACE achieves state-of-the-art latency-fidelity trade-offs. We will release our code soon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.