acceptodds
Under review as a conference paper at ICLR 2027

PhysLDM: Latent Diffusion for High-Fidelity Deformable Simulation

Abstract

Neural simulation of high-fidelity deformable bodies is a foundational challenge in computer graphics and physical AI. However, predicting long-horizon dynamics for 3D volumetric meshes presents multifaceted and tightly coupled challenges. Step-by-step autoregressive methods lack the global temporal context necessary for long-horizon consistency. Conversely, direct multi-frame one-shot prediction is computationally prohibitive at native resolutions. Thus, compressing these dynamics into a spatiotemporal latent space becomes necessary, yet it remains largely unexplored for mesh-based volumetric physics. Furthermore, the optimal predictive paradigm for this task remains an open question. Specifically, it is unclear whether deterministic regression or generative diffusion is better suited for complex physical simulation. To address these coupled challenges, we introduce PhysLDM, a unified latent-diffusion paradigm for one-shot volumetric deformable simulation. The foundation of PhysLDM is a novel holistic spatiotemporal VAE (ST-VAE). In our explorations, we find that standard temporal compression (e.g., video VAEs) loses high-frequency information and causes severe “staircase” artifacts. Our holistic ST-VAE effectively resolves this issue, achieving 2.48 mm precision on meter-scale scenes at up to 78 token compression. Based on this reliable latent space, we systematically compare regression and diffusion paradigms. Our experiments uncover a key modeling insight: complex deformable dynamics are inherently chaotic, and in this regime deterministic regression tends to produce damped, non-physical averages, whereas diffusion is better suited. Consequently, we employ a Latent Diffusion Model (LDM) that naturally respects this stochasticity to generate physically plausible trajectories. By training one single unified model on an Objaverse-scale simulation dataset in a constitutive-model-agnostic manner, PhysLDM achieves strong zero-shot generalization to unseen, out-of-distribution object datasets (GSO, Toys4K). To our knowledge, PhysLDM yields the first high-fidelity spatiotemporal autoencoder and latent-diffusion paradigm for volumetric deformable dynamics, offering a scalable and robust approach to neural simulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.