acceptodds
Under review as a conference paper at ICLR 2027

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

Abstract

Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current Vision-Language-Action (VLA) models are predominantly pretrained on 2D images without explicit 3D geometric supervision, which can limit their ability to capture fine-grained spatial relationships. Prior work partially addresses this limitation by aligning LLM-level visual tokens with features from 3D-aware foundation models. However, our probes show more localized attributions around task-relevant regions at the visual encoder output, along with larger action-prediction errors under encoder-level masking at higher tested masking ratios than under masking at the probed LLM layer. These observations motivate direct geometric supervision of the visual representations supplied to downstream action prediction. We propose VEGA (Visual Encoder Grounding Alignment), a lightweight framework that grounds spatial knowledge directly at the visual encoder output, avoiding the need to select an internal LLM layer for alignment. VEGA uses frozen spatial teachers obtained by fine-tuning visual backbones with multi-view-consistent feature targets rendered from 3D Gaussian representations. A lightweight projector aligns encoder outputs with teacher features through a patch-level cosine distance loss, jointly optimized with the action prediction objective. The teacher and alignment projector are used only during training, leaving the policy architecture and inference cost unchanged. Extensive experiments on RoboTwin 2.0 and real-world manipulation tasks demonstrate that VEGA achieves higher average success rates than existing implicit spatial grounding baselines, with improvements on both OpenVLA-OFT and highlighting its applicability across VLA architectures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.