acceptodds
Under review as a conference paper at ICLR 2027

DynPrompt-Driven: Unleashing Motion Dynamics in Autoregressive Video Generation

Abstract

Autoregressive video generation has emerged as a promising paradigm for efficient streaming and video synthesis, and has witnessed rapid progress in recent years. However, existing autoregressive video diffusion models are prone to motion dynamics degradation during generation: in certain scenarios, videos become increasingly static as generation proceeds, and later segments tend to exhibit weaker motion than early ones. In this work, we ask a central question: how can we explicitly unleash and preserve motion dynamics in autoregressive video generation, thereby enhancing overall video dynamism? Through empirical analysis, we identify a previously under-explored phenomenon: the temporal variation of text cross-attention gradually decreases during autoregressive rollout, as reflected by a progressive reduction in the KL divergence between consecutive text cross-attention maps. Motivated by this observation, we propose DynPrompt-Driven, a dynamic-prompt-driven framework for autoregressive video diffusion. Specifically, we introduce Temporal Prompt Recalibration (TPR), which dynamically recalibrates static prompt embeddings using the accumulated visual KV cache, enabling different textual tokens to be adaptively activated as time evolves. Furthermore, we design Motion-Guided Gradient Alignment (MGGA), which leverages high-motion query videos to measure the dynamic degree of generated videos, steering distribution matching distillation (DMD) toward high-dynamism directions during optimization. Extensive experiments on video generation benchmarks demonstrate that our method substantially improves dynamic degree over strong autoregressive video generation baselines, while introducing only negligible additional parameters and computational overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.