GeoFirst: Learning World-Consistent Geometry for Robust Robot Manipulation
Abstract
World-action models (WAMs) have shown promise for embodied intelligence, but their core value lies in the latent representations learned through prediction rather than explicit future imagination. Existing WAMs rely on RGB video prediction, which lacks 3D geometric structure, while depth-based approaches treat depth as an auxiliary task from limited viewpoints—yielding representations that lack cross-view consistency and degrade under observation perturbations. We present GC-WAM (Geometry Consistent World-Action Model), which replaces RGB video prediction with metric depth prediction as the primary future-state supervision signal, encouraging the video foundation model to learn geometry-grounded latent representations rather than appearance-centric ones. GC-WAM introduces per-view depth normalization to preserve geometric detail across camera views, cross-view world consistency via an InfoNCE-based loss for viewpoint-invariant representations, and a geometry-first curriculum that first learns geometry then jointly learns geometry and action. Experiments on LIBERO Plus show that GC-WAM achieves 79.7% average success rate (+29.8 pp over RGB baseline), and achieves 85.4% on RobotTwin Clean2Clean. Different Views, Same World.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.