acceptodds
Under review as a conference paper at ICLR 2027

From Pixels to Executable Physics: Behavior-Grounded World Models for Embodied Action Consequence Prediction

Abstract

In outdoor environments, varying terrain and contact conditions can cause identical action commands to produce different motion outcomes for quadruped robots. Predicting these outcomes is essential for reliable navigation and informed path selection. Existing video world models can generate visually plausible futures, yet often struggle to capture robot dynamics and terrain interactions, leading to overly optimistic predictions of action outcomes. We present DRWM, a physically consistent world model that incorporates robot–terrain contact dynamics to predict the consequences of quadruped actions. Starting from an initial RGB image, our method reconstructs scene geometry and integrates dense, dimensionless priors over terrain contact properties, calibrated using real-world action–response data. The resulting terrain model supports physics simulation with spatially varying contact properties such as friction. Given the robot’s initial state and candidate action sequences, we simulate its future motion and use the resulting camera trajectories and rendered geometry to guide the training of a video world model. Generated future observations are then used to update the terrain reconstruction, enabling iterative model rollouts without additional real-world images. Extensive experiments demonstrate that DRWM substantially outperforms geometry-only baselines and baselines with uniform physical parameters on key physical metrics characterizing action outcomes, while achieving state-of-the-art performance on all reported video evaluation metrics. These physically grounded predictions support more informed action selection: across 90 challenging scenarios, an 8B vision–language model augmented with DRWM outperforms a 32B vision–language model on five of six evaluation metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.