acceptodds
Under review as a conference paper at ICLR 2027

CacheFlow: Unifying Step Distillation and Feature Caching for Few-Step Generation

Abstract

Step distillation and feature caching accelerate diffusion models by reducing sampling steps and per-step computation, respectively. Under extreme acceleration, distillation struggles to approximate teacher trajectories, sacrificing semantic fidelity and fine details; training-free feature caching degrades on the resulting sparse trajectories, amplifying error propagation when combined with distillation. We introduce CacheFlow, which addresses extreme few-step acceleration from a frequency-aware perspective. We identify asynchronous frequency evolution as the primary source of quality degradation: low-frequency semantics stabilize rapidly, while high-frequency details suffer trajectory collapse and cross-step distortion. We therefore propose frequency-aware policy distillation to provide adaptive multi-band supervision for faithful trajectory approximation and detail preservation. A lightweight, learnable cache policy captures rapid cross-step feature dynamics, overcoming rigid extrapolation in sparse regimes. Jointly optimizing the few-step student and cache policy suppresses caching errors. Experiments on FLUX.1-dev, Qwen-Image, Qwen-Image-Edit, and Qwen-Image-2.1 demonstrate competitive visual quality across text-to-image generation and image editing, with a measured speedup of up to 55.31. Code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.