PhysChain: Constraint-based Physical Chains for Visual State and Transition Generation
Abstract
While existing methods for video generation leveraging the physical reasoning of Vision-Language Models (VLMs) have made impressive progress, they focus on injecting the knowledge into video generation through designated representations (e.g., physical descriptions, property labels, or motion trajectories), bounding the dynamics they can express. In this work, we present PhysChain, a framework that simulates a chain of future states with VLMs to generate physically valid videos from a single image. Our key idea is to simulate the chain of states and transitions under individually verifiable physical constraints to tractably handle real-world complex dynamics, highly entangled and implicit. For consistent and reliable simulations, we design a predict-and-correct loop with sampling-based constraint derivation, and further introduce a novel confidence-adaptive rendering scheme that tolerates the prediction errors that still remain after the simulations. In the experiments, PhysChain consistently outperforms competing physics-aware methods across five physics benchmarks, ranking first on most of them even among large-scale video foundation models, and generalizes to complex real-world dynamics with various potential applications. Videos are available at : https://anonymous-submission-384590248.github.io/anonymous-page.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.