acceptodds
Under review as a conference paper at ICLR 2027

ViDA: Exploiting Temporal Dependencies for Autoregressive Video Distillation

Abstract

Autoregressive video distillation supervises predictions that also serve as context for subsequent generation. Causal self-rollout exposes the student to its own his- tory, but a single re-noising level in Distribution Matching Distillation (DMD) couples the preservation of that context with the correction of generated content. We present ViDA (Video Distillation for Autoregressive generation), which in- corporates temporal dependencies into DMD supervision through ordered score queries. Earlier chunks receive lower, nonzero corruption to retain context while remaining subject to correction. With both score estimators bidirectional and inde- pendent of the generator, the ordered query improves VBench Total by 0.75 over marginal-matched scalar queries, and by 0.64 and 1.40 over shuffled and reversed noise assignments. ViDA also replaces the separate fake-score network with a role-conditioned shared backbone, improving quality at a fixed update ratio. Com- bined with fewer fake updates, it outperforms the independent causal estimator at r= 5 at half the student post-training GPU-hours and 13.6% lower peak allocated memory. On five-second generation, the final 1.3B generator distilled from a 14B teacher reaches 85.62 VBench Total: 1.12 points above matched scalar querying within the same recipe and 1.38 above the locally evaluated official Self Forcing checkpoint. Evaluations on Cosmos and 60-second videos extend these compar- isons across backbones and generation horizons. The autoregressive inference procedure remains unchanged.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.