acceptodds
Under review as a conference paper at ICLR 2027

GeoFirst: Learning World-Consistent Geometry for Robust Robot Manipulation

Abstract

World-action models (WAMs) have shown promise for embodied intelligence, but their core value lies in the latent representations learned through prediction rather than explicit future imagination. Existing WAMs rely on RGB video prediction, which lacks 3D geometric structure, while depth-based approaches treat depth as an auxiliary task from limited viewpoints—yielding representations that lack cross-view consistency and degrade under observation perturbations. We present GC-WAM (Geometry Consistent World-Action Model), which replaces RGB video prediction with metric depth prediction as the primary future-state supervision signal, encouraging the video foundation model to learn geometry-grounded latent representations rather than appearance-centric ones. GC-WAM introduces per-view depth normalization to preserve geometric detail across camera views, cross-view world consistency via an InfoNCE-based loss for viewpoint-invariant representations, and a geometry-first curriculum that first learns geometry then jointly learns geometry and action. Experiments on LIBERO Plus show that GC-WAM achieves 79.7% average success rate (+29.8 pp over RGB baseline), and achieves 85.4% on RobotTwin Clean2Clean. Different Views, Same World.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.