GRAM-VLA:Giving Vision-Language-Action Models a Sense of 3D Space
Abstract
Embodied intelligence is increasingly moving toward more complex manipulation tasks. Compared with single-arm tasks, dual-arm tasks require coordination of the two arms' spatial positions and temporal order. This places greater demands on the model's joint vision-language understanding and spatial perception. Most Vision-Language-Action (VLA) models excel at dual-arm tasks by leveraging large-scale 2D vision-language pre-training, but their reliance on RGB images limits spatial reasoning critical for real-world interaction. Adding 3D information can improve spatial perception. However, without extra large-scale 2D–3D data pre-training, it remains challenging to fully use 3D information for action generation while remaining consistent with pretrained 2D semantics. To address this challenge, we propose GRAM-VLA, a 3D-enhanced dual-arm VLA model without additional large-scale 2D–3D pretraining. GRAM-VLA keeps the pretrained vision-language backbone (VLM) frozen and adds a lightweight depth encoder to form an independent RGB–Depth dual-branch architecture. It establishes token-level spatiotemporal RGB–Depth alignment and alternately injects language, RGB, and depth conditions into the action generation backbone, allowing depth information to fully participate in action generation. Furthermore, we use four-modal GRAM alignment to mitigate multimodal semantic shift caused by the newly introduced depth information. On eight real-robot tasks aligned with the RoboTwin 2.0 simulation protocol, GRAM-VLA achieves an average success rate of 53.6%, outperforming both 2D- and 3D-based models. In controlled experiments, GRAM-VLA outperforms baseline-RDT by 40.0 and 72.5 percentage points under appearance and height shifts, respectively, and by 22.5 percentage points on long-horizon task completion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.