acceptodds
Under review as a conference paper at ICLR 2027

PointmapVLA: Scaling Robot Learning with Robot-Centric Pointmaps

Abstract

Vision-language-action models (VLAs) often struggle with precise manipulation because image observations do not directly provide the metric 3D geometry of the physical environment. We introduce PointmapVLA, the first VLA pretrained at scale with robot-centric pointmaps for joint 3D spatial understanding and robot control. Because pretrained VLM backbones lack a prior for interpreting pointmaps, we jointly pretrain on pointmap-augmented robot demonstrations and 3D spatial VQA, supervising 3D grounding, metric localization, spatial relation reasoning alongside action prediction. To support this pretraining, we construct a pointmap-paired corpus of 76.3M robot demonstration frames and 10.7M 3D spatial VQA examples. Across three simulation benchmarks and three real-world tasks, PointmapVLA delivers strong manipulation performance, establishing a new state of the art with a 92.4% success rate on LIBERO-Plus. In real-world manipulation under unseen target heights, lighting, and camera viewpoints, it achieves 1.8 the success rate of the strongest geometry-aware baseline. We will release the training code, model weights, and pointmap-paired training corpus.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.