Self-Reflecting Video Diffusion with Boosted Motion Dynamics
Abstract
Modern video diffusion models have demonstrated strong performance, enabling the generation of high-fidelity and temporally coherent videos. A key component behind this success is temporal attention, which enables tokens to exchange information across frames to maintain consistency over time. However, we argue that this very mechanism also contributes to limited motion dynamics, a persistent failure mode where generated videos remain overly static. Specifically, we find that weak motion is closely related to tokens over-attending to other frames, a pattern that emerges as early as the initial denoising steps. Building on this observation, we propose Self-Reflecting Optimization (SRO), a training-free framework that optimizes the initial noise to encourage tokens to attend to their own frame, while keeping it close to the original sample and consistent with the Gaussian prior. Since SRO modifies neither the model nor the conditioning signals, it applies to both text-to-video (T2V) and image-to-video (I2V) generation, in stark contrast to existing approaches tied to specific conditioning types. Extensive experiments on both T2V and I2V backbones demonstrate that SRO substantially boosts motion dynamics with only marginal degradation in text-video alignment. Our project page, including additional demo videos, is available at https://sro-psi.vercel.app/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.