AlignFlow: Multi-Axis Cache Scheduling for Efficient Continuous Video Editing
Abstract
Diffusion transformers now dominate video generation, and training-free feature caching is the standard way to cut their denoising cost, yet every existing cache is designed and evaluated for a single generation call. Real editing is not a single call: users issue a sequence of cumulative edits, each conditioned on the previous output. This discards the redundancy that exists across consecutive edits and makes approximation error compound recursively, so single-turn fidelity cannot characterize a session. Extending acceleration across edits is not a matter of adding a second cache. Raw cross-edit reuse is intermittent and, used alone, needs far more exact evaluations than a step-axis cache; finer token-level reuse traded a latency reduction for a memory increase in our diagnostic; and the edited region is precisely where past information is invalid. We present AlignFlow, a training-free controller that predicts the classifier-free-guided output and schedules reuse jointly along the step, region, and edit axes. Damped first-order extrapolation of the guided residual supplies the acceleration backbone; a progress-aligned bank of the previous turn's exact outputs supplies a conservatively mixed cross-edit source, which a dilated latent mask forbids inside the edited region; and a closed-loop calibrator admits each source only when its measured per-region, per-phase reliability clears a threshold, otherwise falling back to an exact step. On five-turn editing chains with VACE Wan2.1-1.3B, AlignFlow reaches a speedup at dB PSNR, dB above the strongest accelerated baseline at lower latency, and transfers to real video and a second architecture. Session-average gains exceed final-turn gains, so we report per-turn fidelity and drift together rather than claiming that recursive drift is removed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.