Geometric Action Model: Repurposing a Geometric Foundation Model for Robot Policy Learning
Abstract
Language-conditioned robot manipulation policies must follow user instructions while reasoning jointly about objects, camera viewpoints, and robot actions in the 3D physical world. Recent vision-language-action models (VLAs) and video world-action models (WAMs) inherit strong semantic and temporal priors from large-scale foundation models. However, these models operate primarily in 2D image space, leaving the 3D geometry essential for contact-rich manipulation only implicitly represented. In this paper, we introduce the Geometric Action Model (GAM), which directly repurposes a pretrained geometric foundation model (GFM) for policy learning. We show that the GFM's rich visual representations and geometric priors provide an effective foundation for manipulation. However, a GFM alone does not natively support future prediction or action generation. GAM bridges this gap by splitting the GFM at an intermediate layer and inserting a causal future predictor between its shallow and deep blocks. The shallow blocks encode observations into geometric features, while the predictor integrates language instructions, proprioception, and action history to forecast future geometric features and action representations. These predictions are then processed by the remaining pretrained GFM blocks to reconstruct future geometry and decode robot actions, with future-feature and future-depth losses explicitly supervising geometric forecasting. This design unifies geometric perception, future prediction, and action decoding within a shared pretrained backbone. Ablations show complementary benefits from geometric initialization and future-prediction supervision for out-of-distribution robustness. Across diverse simulation and real-robot benchmarks, GAM achieves competitive success rates while requiring smaller model size and lower inference latency than representative foundation-model-scale baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.