ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models
Abstract
Action-conditioned world models (ACWMs) have shown strong promise for video prediction and decision-making, yet existing benchmarks are largely restricted to egocentric navigation or narrow robot manipulation, covering only a small part of the physical interactions a general world model must capture. We introduce **ACWM-Phys**, a simulated benchmark spanning four physical regimes (rigid-body, deformable, particle, and kinematic dynamics) across eight environments, each paired with a controlled out-of-distribution (OoD) split that shifts a single physical factor so failures are attributable. Alongside it we release ACWM-DiT, a latent video diffusion transformer baseline. Because every environment is simulated, we score rollouts not only with perceptual metrics but directly against physics: point-track error, conservation of simulator-conserved quantities, and a zero-action counterfactual. Both families of metrics agree on the same trend: OoD degradation is governed by task complexity rather than physics category, with low-dimensional geometric tasks nearly unaffected and contact-rich deformation and high-DoF kinematics degrading sharply. A second, architecturally distinct model post-trained from a pretrained video foundation model shows the same degradation, indicating the gap is not an artifact of our backbone. The physics metrics additionally expose failures that perceptual metrics miss, such as mass that is not conserved under distribution shift. We release ACWM-Phys as a diagnostic instrument for physically grounded world modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.