WSA: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control
Abstract
World-Action Models (WAMs) have emerged as a promising paradigm for robot control by using predicted world evolution to guide action generation. However, existing WAMs largely follow a one-way world-to-action paradigm. Real-world interaction, in contrast, is inherently driven by bidirectional temporal causality with implicit physics: the evolving world informs what actions should be taken, while the executed actions, in turn, determine how the world evolves. Without explicitly modeling this action-to-world direction, WAMs cannot assess whether the consequences of their generated actions align with the world evolution they predict, preventing WAMs from capturing consistent interaction dynamics from video demonstrations and thereby constraining their generalization capabilities in real-world scenarios. In this paper, we introduce WSA, a new modeling paradigm for embodied foundation models. WSA develops a World-Spatial-Action understanding by jointly modeling three interconnected processes: how actions transform the 3D world (world level), how spatial states evolve toward desired outcomes (spatial level), and how target 3D world guide appropriate robot behaviors (action level). By integrating these processes, WSA enables a more coherent and tightly coupled understanding of embodied interactions. Compared with prior WAMs trained on tens of thousands of hours of real-robot data, WSA leverages 6k hours of video demonstrations (only 1k from real robot), demonstrating strong data efficiency. Moreover, even our 3B-parameter WSA surpasses larger WAMs with 6B–8B parameters, highlighting its parameter efficiency. Experiments show that WSA achieves a 93.3% success rate on RoboTwin2.0, delivering an average 20% improvement over state-of-the-art models on real-world OOD robot control tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.