Gaze2Act: Gaze-Guided Vision-Language-Action Policies for Interactive Robot Manipulation
Abstract
Interactive robot manipulation requires specifying which object to manipulate, where to interact with it, and how the target changes during execution. Language instructions alone often leave these aspects of human intent underspecified. We introduce Gaze2Act, a vision-language-action (VLA) framework that incorporates human gaze as a dynamic spatial cue to complement language instructions. Gaze2Act maps first-person gaze into the robot’s view through cross-view semantic matching, yielding an object mask and a gaze point for coarse-to-fine target specification. It integrates these cues into the VLA policy through perception-level prompting and action-level conditioning, enabling precise interactions and target updates during execution. We evaluate Gaze2Act on 16 real-robot tasks spanning seven task categories using a Unitree G1 humanoid. Gaze2Act outperforms the evaluated baselines in intent accuracy and task success rate, with improvements in object disambiguation, fine-grained interaction, and dynamic intent steering. These results support human gaze as a complementary input for specifying spatial intent and guiding VLA policies during interactive manipulation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.