MM-World: A Unified Video World Model for Egocentric Navigation and Manipulation
Abstract
Action-conditioned video world models enable users to explore and interact with virtual worlds through their own actions. However, existing models typically focus on either navigation or local manipulation, leaving their combination poorly supported, even though most real-world tasks require moving and manipulating in tandem. To address this, we present MM-World, a unified video world model for egocentric navigation and manipulation controlled by head and hand motion. Our key insight is to express these heterogeneous action-conditioning signals in a shared image space. We reproject scene geometry reconstructed from the initial image along the head trajectory and overlay hands mesh rendered in their prescribed poses. The resulting visual condition reuses the pretrained model's native visual-conditioning pathway without action-specific input modules, while providing a shared interface for joint training on navigation and manipulation videos. We further adopt an autoregressive generation scheme to support long-horizon interactions. Moreover, to evaluate these capabilities comprehensively, we curate a unified dataset and establish a four-track benchmark with 846 test cases spanning navigation, manipulation, mobile manipulation, and long-horizon generation. Extensive experiments show that MM-World outperforms previous methods by a large margin across visual quality and action-following.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.