-VLA: Visual Self-Grounding for Vision-Language-Action Policies
Abstract
Vision-language-action (VLA) policies have shown strong promise as generalist controllers for robot manipulation, yet their action predictions can remain fragile under changes in camera extrinsics. We identify a systematic failure pattern in which camera perturbations induce directional drift in predicted robot trajectories, revealing that current VLAs may exploit image-space shortcuts, such as pixel-coordinate correlations, rather than grounding actions in robot–object geometry. To bridge this gap, we introduce -VLA, a Self-Supervised alignment approach to incentivize visual Self-grounding capability in VLA policies. Our central insight is that robot motion provides intrinsic spatial supervision for localizing the robot itself. To implicitly guide the policy to attend to the robot, we formulate an auxiliary grounding objective to recover motion-derived localization targets from its visual representations. Across simulation benchmarks and real-world manipulation tasks, -VLA consistently improves average task success on two leading VLA backbones, and GR00T-N1.6. Under camera-extrinsic perturbations, -VLA increases the average success rate on real-world tasks from 16.4% to 29.4%. These findings highlight the importance of visual self-grounding for robust vision-language-action policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.