GeoMem-VLA: Internalizing 3D Scene Understanding into Temporal Memory for Vision-Language-Action Models
Abstract
Geometric structure provides a direct description of the 3D world and is essential for robotic manipulation and world understanding. Most VLA models primarily rely on 2D representations of the current observation, while geometry-aware approaches often require additional 3D geometric information. However, such geometric information, typically inferred independently from individual observations, is often discarded rather than accumulated into a persistent task representation. Our key insight is that action generation does not require a complete, renderable 3D reconstruction at every step; instead, a policy needs a compact, task-relevant geometric abstraction that can be continuously updated over time. Based on this insight, we introduce GeoMem-VLA, a VLA framework that internalizes explicit 3D supervision into dedicated geometric tokens, without directly aligning the VLM's native visual embeddings to the teacher feature space, and maintains them as temporal memory for action generation. During training, a frozen geometry foundation model supervises learnable GeoPlan tokens through a geometric flow-matching objective over the current scene point map and future end-effector tracks. These tokens are subsequently integrated with the memory of observation representations, where historical information is retrieved, adaptively fused with current observations, and consolidated to keep event frames as the task progresses. At inference, the policy retrieves historical representations and fuses them with the current working representation to predict actions, without invoking the 3D teacher or explicitly reconstructing the scene. After benchmark-specific downstream adaptation, GeoMem-VLA achieves 98.9% average success on LIBERO and 91.10% / 91.24% on the clean/randomized settings of the RoboTwin benchmark. On the real-world tasks, GeoMem-VLA achieves 78.0% success score. These results indicate that internalizing and remembering 3D scene understanding is a promising alternative to repeatedly reconstructing explicit geometry for robotic control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.