ASTAR: Accelerating Spacetime Autoregressive Video Generation via Appearance Inheritance
Abstract
Spacetime autoregressive (STAR) models extend next-scale visual autoregression to video through coarse-to-fine spacetime prediction, yet their late high-resolution scales remain computationally expensive. We analyze how representations evolve during clip-wise generation and uncover Appearance Inheritance: much of the visual state established by the first clip persists as generation proceeds, while larger changes concentrate around evolving content. Building on this observation, we introduce ASTAR, a training-free acceleration framework that uses the first clip as a transformation reference for later clips. ASTAR reserves full Transformer computation for selected tokens and approximates the remaining positions with first-clip-derived layer transformations, while residual evidence from current computation constrains the approximation before multi-scale accumulation. ASTAR achieves 2.40× and 1.93× end-to-end speedup at 720p and 480p, respectively, with less than a 1% relative decrease in VBench Total at either resolution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.