PhysPilot: Top-Down Reinforcement Learning of Physical Priors for Video Generation
Abstract
Recent video generation models are capable of generating visually realistic videos, but often fail to adhere to physical laws, limiting their ability to generate physically plausible videos and serve as "world models". To address this issue, we propose PhysPilot, a generalizable physical law learning framework designed to enhance the physical plausibility of video generation. Specifically, PhysPilot is based on the image-to-video task where the model is expected to predict physically plausible dynamics from the input image. Since the input image provides physical priors like positions, materials, and interactions of objects in the scenario, we devise a physical encoder, PhysEncoder, to extract such physics-relevant visual cues as an extra condition. However, how to help the model learn to utilize these physical priors to assist generation remains an open question. Without a well-established definition for physical representation, we can not conduct straightforward supervision of PhysEncoder for training. To solve this challenge, we adopt a top-down optimization strategy, where the PhysEncoder and the video generation model are jointly optimized based on the physical plausibility of the final generated videos using reinforcement learning (RL). Through this top-down optimization, the PhysEncoder can effectively capture implicit physical cues from the starting image, and the video model learns to better utilize them, thereby implicitly learning physical laws. Experiment results prove that our method significantly enhances the model’s physics-awareness as a plug-in, and demonstrates strong generalization on both specialized proxy tasks and diverse open-world physical scenarios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.