GeoWAM: Geometry Gradient Transfer For World Action Models
Abstract
World action models learn robot policies through video prediction and action generation, but predicting visual appearance does not explicitly encourage representations of how scene geometry evolves. For manipulation, the critical question is not only what a future scene will look like, but how the spatial relationships between the robot and surrounding objects will change. We argue that predicting geometry trajectories provides complementary supervision for understanding how scenes evolve during interaction and learning policies that generalize beyond visual appearance. We introduce GeoWAM, which augments video–action learning with generative modeling of geometric evolution. A frozen 3D foundation model and a spatiotemporal geometry VAE transform multi-camera video windows into structured geometry latent sequences, which a geometry expert learns to predict through flow matching. Our central mechanism, geometry gradient transfer, uses asymmetric attention: the geometry expert reads video and action representations, while neither expert reads geometry. The geometry objective thus backpropagates into the deployed experts, encouraging their representations to support geometric prediction without making action generation depend on geometry inputs. The geometry branch is omitted at deployment, supporting both joint video–action denoising and action generation without future-video rollout. On LIBERO, no-rollout GeoWAM achieves 98.3% success, improving over a matched baseline by 2.3 percentage points, and attains 65.2% on LIBERO-Plus. Our joint variant reaches 76.49% category-averaged success on LIBERO-Plus. These results support modeling geometric evolution as a training-time source of policy improvement without test-time geometry inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.