acceptodds
Under review as a conference paper at ICLR 2027

Speculative Decoding for Autoregressive Video Generation

Abstract

Autoregressive video diffusion has emerged as a promising paradigm for streaming video synthesis, yet accelerating inference remains a critical challenge. While speculative decoding is the dominant acceleration strategy for large language models (LLMs), adapting it to video generation remains non-trivial, as continuous spatiotemporal video tensors lack discrete token probability distributions for exact rejection sampling. To bridge this gap, we introduce SDVG, a training-free framework that brings speculative decoding to autoregressive video diffusion via reward-guided block verification. Instead of discrete logits, SDVG employs a lightweight image-quality router with worst-frame score aggregation to evaluate candidate blocks proposed by a compact draft model. Accepted drafts are directly committed to the target model's key-value (KV) cache, while rejected blocks trigger target regeneration. By enforcing mandatory target-based generation on the initial anchor block to fix global scene composition, a single calibrated threshold provides a simple, continuous control knob over the quality–speed trade-off. We further introduce SDVG-hybrid, which integrates step-level trajectory stitching on rejected blocks to reuse draft compute and incorporates an on-demand temporal cascade to minimize verification overhead. Evaluated on 1,003 MovieGenVideoBench prompts (832 × 480), standard SDVG achieves a 1.59× speedup while preserving 98.1% of target visual quality (VisionReward), reaching up to 2.05× acceleration at 95.9% retention (+17.4% over draft-only generation). Furthermore, SDVG-hybrid advances the Pareto frontier to a 1.88× speedup at 98.0% quality retention and 1.96× at 96.9% retention. Requiring no retraining or architectural modifications, SDVG serves as a practical, plug-and-play accelerator for modern autoregressive video pipelines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.