EVGT: Embodied Visual Geometry Transformer
Abstract
3D reconstruction is fundamental to robotic perception, providing the spatial structure needed to understand scenes and object motion. Despite recent progress in visual geometry foundation models, a unified geometry foundation model tailored to robotic scenes remains lacking. To address this gap, we propose an Embodied Visual Geometry Transformer (EVGT), a metric-scale geometry foundation model for robot video that integrates multi-view observations to recover scene geometry with the motion and 6D pose of the object of interest. From RGB observations, EVGT predicts camera poses, metric depth, point maps, and dense 3D tracks together with 6D pose, size, oriented box, mask, and visibility of the object across time in a shared reference frame. We extract visual tokens from RGB observations and augment them with object queries, with optional prompt tokens for target specification. A spatiotemporal transformer models interactions among these tokens across views and frames, and multiple prediction heads decode the resulting features into scene geometry, 3D motion, and object states. We adopt a two-stage training strategy that combines heterogeneous scene-geometry supervision with partially labeled object-state data. We report point-map, depth, and camera-pose results on six robot-operation sources, dynamic 3D tracking on four core sources, and object-center and rotation diagnostics on six object sources. In downstream closed-loop manipulation, augmented with frozen EVGT features reaches 86.31% success on LIBERO-Plus and 59.17% on RoboCasa24, compared with 84.67% for the LIBERO-Plus starting policy and 55.1% for a RoboCasa-fine-tuned RGB-only baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.