acceptodds
Under review as a conference paper at ICLR 2027

Physics-Guided Video Prediction via Learned Dynamics from Video Observations

Abstract

Future video prediction for physical systems must preserve both visual realism and physical dynamic evolutions. Existing video generators often model motion non-physically or rely on externally user-defined manipulations, while methods that infer physical dynamics from video typically focus on identifying parameters or trajectory prediction, with limited integration into future video generation. This leaves a research gap between explicit physical prediction and physically grounded future video generation. We introduce LPhyVideo ("L" stands for "Learned"), a physics-guided video prediction framework via learned physical dynamics from the video, within a video-to-physics-to-video formulation for systems whose low-dimensional physical states are reconstructed from visual observations. LPhyVideo maps tracked key pixel locations to physically interpretable states, and rolls out via a learned state-space dynamics model; subsequently, LPhyVideo reconstructs the predicted states back into image space, and with the aid of a reference frame, converted into a video and edit mask for a constrained refinement of final generation. Across six diverse physical systems, LPhyVideo achieves the highest average Spatial IoU, Spatiotemporal IoU, Weighted Spatial IoU and the lowest MSE and Trajectory Error among comparisons.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.