acceptodds
Under review as a conference paper at ICLR 2027

GeoWRA: Geometric Latent World Residual Action Models

Abstract

World models and vision-language-action (VLA) models provide complementary capabilities for robotic manipulation through predictive modeling and multimodal action generation. However, effectively integrating world models and VLAs into a unified architecture remains a key challenge, as existing approaches are structurally decoupled, lacking intrinsic interaction between world prediction and action generation. To address this challenge, we propose Geometric Latent World Residual Action Models (GeoWRA), a unified framework that couples latent world prediction with action generation. First, we establish a bidirectionally coupled architecture in which the VLA's latent actions condition future prediction, while predicted futures in turn guide action generation through bounded world residual feedback. Second, to capture the relational structure of embodied scenarios, we introduce a Euclidean-hyperbolic product representation that jointly models future features and their relationships by embedding predicted future states in hyperbolic space, providing a geometric inductive bias for learning task-relevant predictive representations. Third, we construct a joint action and geometric world objective that integrates action supervision with future prediction under hyperbolic geometric regularization, enabling geometrically consistent future representations to guide action generation. Extensive evaluations in both simulation and real-world settings show that GeoWRA delivers competitive performance across a diverse range of robotic manipulation tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.