World Before Pixels: Agentic Physical Simulation for Text-to-Video Synthesis
Abstract
Generative text-to-video synthesis has achieved striking visual realism, yet current models still struggle to ensure that scene geometry supports the requested interaction and that the resulting motion follows physical laws. Faithfully modeling the world requires understanding not only visual appearance, but also object geometry and material properties, physical states and interactions, and the causal dynamics that govern how these states evolve—knowledge that is difficult to acquire implicitly from even massive collections of video data. We present an agentic framework that addresses these challenges by constructing, validating, and simulating an explicit physical world before synthesizing its appearance. Given a text prompt, our agent constructs a structured and executable scene whose geometry, materials, physical roles, and simulation parameters expose the variables governing the requested event. Simulation then turns these coupled choices into observable consequences, including rigid-body contact, deformation, and fluid motion, enabling a vision-language controller to assess the simulated outcome and revise the scene through validated edits and re-execution. This closed loop progressively aligns the constructed world with the intended event. Finally, as the requests from the initial text prompt are met, verified geometric priors are rendered from the simulated rollouts, supplying aligned depth, contour, and material controls for a pretrained video model. This combines simulated dynamics with generative appearance synthesis without additional model training. Experiments across diverse physical scenarios show improved physical fidelity over direct text-to-video generation and ablated variants.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.