SpatialWAM: Spatially Grounded World Action Models via Explicit 3D Geometric Learning
Abstract
World action models inherit spatio-temporal priors from video models, but visual plausibility does not ensure geometric accuracy for fine-grained manipulation. We present **SpatialWAM**, which learns spatially grounded video and spatial representations through explicit 3D geometric learning for action generation. During video–spatial pre-training, a spatial expert extracts features from current and future video tokens using spatial and dynamics queries. Spatio-Temporal Gaussian Modeling (STGM) then fuses these features to predict current and future 3D Gaussian scenes shared across views. Multi-view depth supervision in turn shapes both experts’ representations through bidirectional attention, using calibrated RGB-D videos without action annotations. During post-training, a randomly initialized action expert accesses the learned representations through joint attention, while all three experts are fully updated with continued video and depth supervision. At inference, all three experts remain active, while STGM is omitted. SpatialWAM achieves 73.4% success on RoboCasa, 97.8% on LIBERO, and 85.9% on LIBERO-Plus, with 72.0% average success across three real-world tasks. Comparisons using identical pre-training data show gains over video-only and 2D depth pre-training, supporting geometric learning and shared 3D scene modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.