ARD: Agentic Autoregressive Diffusion for Long Video Consistency
Abstract
While diffusion models synthesize high-fidelity short clips, transforming them into consistent and coherent long videos remains challenging. Existing agentic approaches plan and coordinate segment generation, but often leave each segment unchanged once generated, leading to persistent semantic drift over long horizons. We introduce Agentic Auto-Regressive Diffusion (ARD), a training-free agentic architecture that decouples creative synthesis from consistency enforcement. ARD formulates long video synthesis as a closed-loop autoregressive process that synthesizes and self-improves video segment-by-segment against the evolving video world through a Retrieve–Synthesize–Refine–Update cycle. Its architecture combines (i) Multimodal Video Memory, which tracks visual and narrative states across segments; (ii) Adaptive Segment Generation, which selects extrapolation or interpolation according to the narrative transition; and (iii) Hierarchical Test-Time Self-Improvement, which refines boundary frames and video segments at their respective levels. We further introduce LVBENCH-C, a challenging benchmark with non-linear entity and environment transitions to stress-test long-horizon consistency. Across public and LVBENCH-C benchmarks spanning one- to ten-minute videos, ARD outperforms state-of-the-art baselines by up to 30% in consistency and 20% in narrative coherence. Human evaluations corroborate these gains while also highlighting notable improvements in motion and transition smoothness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.