What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
Abstract
Video generators synthesize visually compelling clips, yet often violate basic physical laws–objects float or collide inconsistently, accelerations drift–revealing a gap between visual and physical realism. We propose NewtonRewards, the first physics-grounded post-training framework for video generation based on verifiable rewards. Instead of relying on human or VLM feedback, NewtonRewards extracts measurable proxies from generated videos using frozen utility models: optical flow serves as a proxy for velocity, while high-level visual features serve as a proxy for physical properties. These proxies enable explicit enforcement of Newtonian structure through two complementary rewards: a Newtonian kinematic constraint enforcing constant acceleration dynamics, and a physical properties conservation reward preventing trivial, degenerate solutions. We evaluate NewtonRewards on five Newtonian Motion Primitives (free fall, horizontal/parabolic throw, and ramp sliding down/up) using OpenSora and Wan 2.2 5B on our newly constructed large-scale benchmark, NewtonBench-60K. Across both models in visual and physics metrics, NewtonRewards consistently improves physical plausibility, motion smoothness, and temporal coherence over prior post-training methods. It further maintains strong performance across all primitives and under out-of-distribution shifts in height, speed, and friction. Our results show that physics-grounded verifiable rewards offer a scalable path toward physics-aware video generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.