acceptodds
Under review as a conference paper at ICLR 2027

Render the Robot, Generate the World: One Video World Model Across Embodiments

Abstract

Predicting the consequences of an agent's actions requires modeling both the embodiment's motion and the environment's response. Existing video world models typically learn both jointly from numerical action vectors, even though known robot geometry and kinematics already provide much of the information needed to specify the intended robot motion. A general-purpose world model should capture these dynamics across multiple embodiments within a single model. Our key idea is to render the robot along a proposed joint-space trajectory and condition the video model directly on these future robot configurations. To this end, we introduce DOPPEL, an embodiment-agnostic video world model for learning a dynamics model shared across embodiments. DOPPEL jointly conditions on camera-aligned robot renderings and numerical action trajectories. Using forward kinematics and differentiable rendering, we explicitly provide the model with the robot's intended motion. This allows the shared model to focus its capacity on object motion and interaction-driven changes in the environment. Representing robot motion in image space also provides a common conditioning interface across embodiments with different geometries and action spaces. To enable this approach at scale, we recover missing camera extrinsics from offline robot datasets and align the resulting renderings with recorded observations. Using these aligned renderings, we train a single shared video backbone and conditioning adapter jointly on DROID, RoboMIND, and Bridge, spanning Franka, UR5e, and WidowX embodiments. Across three embodiments, DOPPEL improves PSNR over action-only baselines by 1.31 dB on unseen trajectories and 1.36 dB under OOD evaluation, while predicting more consistent object motion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.