acceptodds
Under review as a conference paper at ICLR 2027

Sample Midway, Condition Deeper: One-Pass Autoregressive Video Generation

Abstract

Discrete-token video generators sample a clip as a grid of thousands of codes, and raster decoding runs one full transformer pass per code (174 s per UCF101 clip on an A100). Parallel decoders sample a group of codes per pass, but the codes in a group are drawn independently, so a choice that the condition leaves open, such as the phase of a pedaling motion, can be drawn differently at different positions. Before a code is drawn, every hidden state is a deterministic function of the condition, so only the drawn code itself can tell another position which option was taken, and this code can re-enter the network at the layer where it was drawn. We propose IntraAR, which samples successive token groups at successively deeper layers of one forward pass. After a group is sampled at layer , the embedding of each drawn code is added to its hidden state, the row is frozen as key–value memory, and the remaining layers predict the next groups conditioned on it. One pass thus performs sampling steps with an exact likelihood. We show that sampling a group earlier pays off when the information it passes to later groups exceeds the fitting error caused by its shallower prediction. On five video benchmarks, IntraAR lowers batch-1 latency by 24.5–41.5× relative to matched raster decoding while I3D-FVD changes by −2.26 to +1.83; on three ImageNet backbones, latency falls by 11.2–29.8× at FID increases of 0.14–0.28. Removing only the embedding of the drawn code, with identical groups and computation, raises UCF101 I3D-FVD from 195.6 to 269.6.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.