Dual-Sim: Interlocked Physical-Generative World Simulator for Robotic Manipulation
Abstract
Action-conditioned video world models, \ie, world simulators, produce the future video of a robot executing given actions, offering a photorealistic environment for robotic manipulation. However, three limitations remain in current literature. (1) The generated robot motion gradually drifts from the given actions as the horizon grows. (2) Every new embodiment requires a separately fine-tuned model. (3) The proprioception state that a policy needs for training and evaluation is missing or inaccurate. In contrast, a physics simulator can compensate for these limitations, since it computes robot motion directly. Therefore, we devise Dual-Sim, a world simulator that interlocks a physics simulator and a generative simulator, combining the kinematic fidelity of the physics simulator with the visual fidelity of the generative simulator. Specifically, Dual-Sim casts world prediction as a conditional transport from the robot kinematics to the video and proprioception state. This transport is learned by a video-state Mixture-of-Transformers, whose video expert and state expert exchange features along one shared flow and predict the video and the state jointly. We further propose kinematic seeding, which replaces the Gaussian source of the flow with an inversion of the rendered robot kinematics. The generative simulator thus refines the robot and completes the scene rather than imaging the future from scratch. Dual-Sim surpasses existing methods in action following, cross-embodiment generalization, and proprioception state accuracy. Closed-loop experiments further establish it as a faithful world simulator.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.