PhysWonder: Building Simulation-Ready Worlds Where Deformation Dynamics Reveal Constitutive Response
Abstract
Recent video generation models show growing potential as world simulators, but physically controlling material-dependent dynamics remains difficult. The same interaction can produce different dynamics depending on how materials deform and whether they recover, reflecting their distinct constitutive responses. Existing approaches guide video generation with simulated motion, forces, or material attributes, but leave this constitutive response implicit. Our key insight is to reveal constitutive response by jointly modeling video generation and physically aligned, appearance-decoupled deformation dynamics. We present ***PhysWonder***, combining two novel modules: ***PhysWonder-3D*** for scalable, physically aligned synthesis and ***PhysWonder-Ctrl*** for learning material control. PhysWonder-3D builds simulation-ready particle worlds from single images and uses a hybrid Material Point Method (MPM) to vary constitutive settings under fixed world and interaction conditions. This produces matched rigid, elastic, sand, and snow dynamics, with RGB videos and deformation-response supervision co-rasterized from the same simulator states. PhysWonder-Ctrl jointly trains a Control Latent Predictor and video generator with response-latent and video supervision. We condition the video generator with the predicted latents, enabling material-controllable generation from a reference image and a physical prompt without simulator rollouts at inference. PhysWonder-3D preserves simulated trajectories while achieving a 2.46× end-to-end speedup over RealWonder in single-image video synthesis. PhysWonder-Ctrl improves material controllability and motion similarity under both explicit and predicted guidance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.