PhysVerse: Deformable Simulation-Ready 4D World Modeling from Monocular Videos.
Abstract
World modeling aims to predict future observations while retaining a scene state for downstream use. Existing reconstruction and generation-centric methods mainly capture geometry and visual dynamics, and rarely represent the spatially varying material parameters a soft-body simulator requires, so the retained state cannot be used directly for deformable simulation. We present PhysVerse, a deformable simulation-ready 4D world model that links reconstruction, video generation, and deformable simulation through a shared scene state. Given monocular video, a feed-forward reconstructor jointly predicts depth, dynamic Gaussian primitives, and dense maps of Young's modulus, Poisson's ratio, and density in a single pass. The reconstructed geometry and material maps initialize MPM-based deformable simulation, while the same depth and material maps condition future video prediction. Generated frames can be passed once again through the reconstructor to recover geometry and material maps. We evaluate PhysVerse on both surgical and general robotic video benchmarks. Experiments show that the reconstructed state supports deformable rollout under MPM, and that depth and material conditioning improve video fidelity and temporal consistency. Our demo is available at: https://anonymous.4open.science/w/doaneqzp-0DBD/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.