TARGET-VLA: Mask-Guided Visual Retrieval for Robotic Manipulation
Abstract
Vision-language-action (VLA) policies benefit from pretrained vision-language representations, but precise manipulation also requires fine-grained object-level and geometric information. However, our controlled comparison indicates that incorporating complementary representations from specialized vision foundation models without explicit spatial guidance yields only modest gains. Motivated by this observation, we introduce TARGET-VLA, a method for mask-guided retrieval of visual representations from frozen segmentation and depth models. It separates the roles of spatial guidance and visual content: prompt-conditioned segmentation masks guide where to retrieve information, while pretrained representations supply the features for action generation. Experiments in simulation and on a physical robot demonstrate improved manipulation performance over retrieval without mask guidance. TARGET-VLA achieves competitive results across simulation benchmarks, including zero-shot transfer to LIBERO-Plus, with real-world gains extending to shifted object positions and object categories absent from the robot demonstrations. These findings demonstrate the value of spatial guidance in integrating semantic and spatial representations to improve robotic manipulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.