CausalMotion: Structured Physical Reasoning as Keyframe and Trajectory Guidance for Training-Free Video Generation
Abstract
Recent advances in diffusion-based video generation have significantly improved visual quality and short-term temporal coherence. However, existing video diffusion models, while effective at modeling local structures and motion, often struggle with long-horizon physical inference, resulting in dynamics that lack physical consistency and causal plausibility. We therefore propose **CausalMotion**, a training-free framework that injects explicit physical reasoning into video generation through structured intermediate representations. Our key idea is to decouple reasoning from generation by leveraging a vision-language model to translate the text prompt into a sequence of causally consistent scene states represented by keyframes, together with object-centric motion trajectories. These representations provide complementary spatial and temporal guidance, which we align and integrate as soft constraints into a pretrained video diffusion model during inference. By explicitly planning object states and dynamics before generation, CausalMotion improves physical plausibility and temporal coherence without additional training or supervision. Extensive experiments demonstrate consistent improvements over existing baselines, particularly in dynamics-intensive scenarios, while maintaining high perceptual video quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.