acceptodds
Under review as a conference paper at ICLR 2027

WAM4D: Learning to Act through Geometry-Grounded World Prediction

Abstract

World-action models couple future visual prediction with robot action generation, yet visually plausible futures do not necessarily capture the spatial changes required for reliable manipulation. We introduce WAM4D, a geometry-grounded world-action model that connects future geometric prediction to robot control while retaining RGB-only sensing. Guided by a pretrained geometric decoder and future-depth supervision, spatial register tokens learn temporally structured features from RGB observation history within a shared video-action backbone. Asymmetric causal attention makes these geometry-supervised features directly available to action queries alongside RGB features, integrating geometric prediction and action generation into single-stage training without future-target leakage. At inference, history and register features are computed once per decision step and cached for flow-based action denoising, requiring neither future-video generation nor dense depth decoding. In controlled experiments on four RoboTwin 2.0 tasks, WAM4D achieves higher average success than an RGB-only WAM and a 4D-WAM reimplementation across their common evaluation conditions, with the largest gains under texture and dynamic lighting changes. Frozen-feature probes further show that its registers retain action-relevant information and are less sensitive to appearance perturbations than RGB-only WAM features. Complementary evaluations with an earlier register-block variant show competitive full-task RoboTwin performance, improved geometric prediction on RT-1 and Bridge-Data, and higher average sub-action success in real-world manipulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.