InternW0-: An Embodied World Model Bridging Predictive Dynamics and Actions
Abstract
World Action Models (WAMs) have emerged as a promising paradigm for generalist robot manipulation by jointly modeling visual dynamics and action generation. A central challenge is how to effectively integrate complementary priors from large-scale pretrained models—including visual dynamics, scene semantics, and geometric and motion understanding—into a unified framework. To this end, we introduce InternW0-, a unified World Action Model that brings together pretrained visual dynamics, scene-level semantic understanding, 4D geometric and motion priors, and action generation within a directed Mixture-of-Transformers framework. Within the World–Action MoT, a pretrained video expert and an action expert interact under scene-grounded semantic guidance from a frozen VLM, while a pretrained 4D foundation model injects geometric and motion priors through training-only distillation. To further translate predictive visual dynamics into representations directly useful for action prediction, we introduce Causal Imprint, which learns future-relevant scene changes from training-only future supervision and makes these predictive representations directly available to the action expert without requiring future-video rollout at inference. To support large-scale joint training, we construct a heterogeneous corpus spanning robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data. We carefully curate and filter these diverse sources, resulting in over 20K hours of processed training data. We pretrain InternW0- on this heterogeneous corpus and demonstrate strong performance across diverse simulation benchmarks and real-robot platforms. We will open-source the model code and weights, infrastructure, and data-processing pipeline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.