acceptodds
Under review as a conference paper at ICLR 2027

Geometric Latent States for Spatially Aware Robot Manipulation

Abstract

Predicting physical scene dynamics under interaction is fundamental to reliable robotic manipulation. World Action Models (WAMs) have emerged as a promising paradigm by framing scene evolution as future video prediction. However, existing WAMs predominantly operate in appearance-oriented video latent spaces, lacking the 3D spatial understanding required for precise manipulation. While per-frame depth prediction constrains distances to visible surfaces without capturing spatial correspondences across views and time, emerging 3D foundation models offer rich spatial representations whose high dimensionality and redundancy make direct prediction computationally demanding. We present GeoAct, a world action model that grounds robot manipulation in latent 3D geometric prediction. GeoAct jointly models complementary future states: RGB latents to capture fine-grained visual dynamics and geometric latents to model evolving 3D scene structure. To obtain a compact spatial state, we design GeoAE to compress multi-layer VGGT-Omega features into a low-dimensional spatiotemporal latent space through joint feature reconstruction and explicit 3D geometric decoding. During training, GeoAct co-denoises future RGB, geometry, and actions, while retaining RGB-to-action inference without online geometry encoding at deployment. GeoAct achieves mean success rates of 84.6% on LIBERO-Plus and 92.6% on RoboTwin 2.0, with strong performance under camera viewpoint perturbations. Real-world experiments further demonstrate GeoAct's improved performance on manipulation tasks requiring precise spatial coordination.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.