acceptodds
Under review as a conference paper at ICLR 2027

BlockDelta: Block Residual Reconstruction for Frame-Sparse Few-Step Video Diffusion

Abstract

Few-step distillation accelerates video diffusion models, yet each remaining step is still computationally expensive, and on 4-step models, cross-step feature reuse yields limited benefits. Frame-sparse inference alleviates this cost by computing only selected frames (anchors) within each transformer block and reconstructing the skipped frames from them. Its effectiveness hinges on two design choices: (1) what to reconstruct for skipped frames, and (2) how to estimate this target from anchors. We propose BlockDelta, a training-free frame-sparse method that systematically addresses both. For (1), we reconstruct the block residual, the difference between a block's output and input, rather than the full output. Reconstructing the full output implicitly re-estimates the already-known input, introducing unnecessary approximation error. Theoretical analysis and empirical evaluation confirm that residual reconstruction reduces block-wise error and boosts generation quality. For (2), we find that lower local reconstruction error does not necessarily yield higher fidelity to dense inference, i.e., sampling without acceleration. Averaging all anchors has the lowest local error but over-smooths inter-frame differences and leads to larger deviation from dense inference, whereas linear interpolation between the two nearest enclosing anchors preserves temporal structure more faithfully, and BlockDelta uses it. On distilled Wan2.2 and HunyuanVideo-1.5, BlockDelta reduces per-step computation by nearly half and keeps the VBench score within 1.1 points of dense inference, while part of the motion is lost, most on Wan2.2 T2V. It also applies to recent audio–video diffusion models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.