acceptodds
Under review as a conference paper at ICLR 2027

Towards Realistic Temporal Dynamics in Video Diffusion Models

Abstract

Text-to-video diffusion models have achieved remarkable progress in visual fidelity and semantic alignment, but still struggle to reproduce the temporal dynamics observed in real videos. We characterize this discrepancy along two complementary dimensions: motion magnitude, which measures how much temporal change occurs, and temporal pacing, which describes how that change unfolds over time. Across three video generators, we observe similar temporal miscalibration patterns, while the required calibration varies across models and individual samples. Based on these observations, we introduce Adaptive Temporal Coordinate Calibration, an inference-time method that reparameterizes the temporal RoPE coordinates of a pretrained video diffusion model. During denoising, we predict temporal statistics from intermediate representations and calibrate them toward real video dynamics. We evaluate our method on LTX-Video-13B, Wan2.1-T2V-1.3B, and CogVideoX-5B using held-out SSv2 and Panda-70M videos with external temporal benchmarks, including ChronoMagic-Bench and Ref4D-VideoBench. On the held-out dataset, our method reduces -Gap by 17.2% and Prog- by 19.0% on average across the three evaluated models, while largely preserving visual quality. The method requires no additional denoising steps and adds less than 0.7% overhead across the evaluated models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.