PhysON: Physicalizing Everything Everywhere All at Once
Abstract
Recovering spatially varying material behavior from video remains challenging. Existing approaches often assume a homogeneous material or restrict inference to a predefined constitutive family. We study visual system identification of spatially heterogeneous scenes, encompassing material variation within objects and across interacting objects. The central challenge is to infer local constitutive differences from the coupled motion of the scene. Our key insight is that a shared constitutive representation can constrain these local responses while allowing their spatial distribution to be inferred from motion. We introduce PhysON, which learns a latent-conditioned constitutive model from analytic material responses and represents scene materials as a continuous spatial latent field. Reconstructed continuum trajectories guide field optimization through differentiable simulation, with the shared constitutive model kept fixed and no object-wise material assignments required. We also introduce a synthetic multiview benchmark covering heterogeneous objects and interactions among different materials. Experiments on public benchmarks and our dataset demonstrate improved motion reconstruction and future prediction across diverse material behaviors, including spatially heterogeneous scenes. Ablations highlight the importance of material-field initialization and constitutive supervision for recovering local material responses from scene dynamics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.