BlockStep: Accelerating Autoregressive Video Diffusion via Online Multi-Model Routing Across Blocks and Steps
Abstract
Autoregressive video diffusion has become a compelling paradigm for streaming video synthesis, with step distillation as its standard path to fast inference. Yet step distillation reduces only the number of steps, not the size of the model run at each one: every step still invokes the full-size network – tens of billions of parameters – and replacing it with a smaller model can yield significant quality degradation. We introduce BlockStep, an online multi-model router: for each prompt, it routes every block and denoising step to a model of the appropriate size – the full-size base model or a smaller model distilled to the same few-step regime – invoking the base model only where it is needed to maintain quality while maximizing speedup. Unlike the leading autoregressive router, which scores each block with an external image-reward model – a costly extra network and an imperfect proxy for video quality – BlockStep does not rely on an extra reward model and is training-free. Specifically, BlockStep consists of three components: (1) a fixed block-0 anchor runs the base model on the first block, recovering most of the quality gap to the smaller model; (2) an online, block-level router upgrades selected blocks to the base model, based on the smaller model's attention entropy; and (3) within each upgraded block, a step-level schedule is selected based on block position. On Krea Realtime 14B with Wan2.1-1.3B Self-Forcing as the smaller model, BlockStep matches the base model's VBench total score at 2.37x wall-clock speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.