Learning Grounded Reasoning from Human Gaze
Abstract
Human visual reasoning is driven by systematic exploration and extraction of task-relevant visual evidence. Human gaze provides a window into this process, making gaze a useful source of guidance for visually-grounded reasoning in vision-language models (VLMs). However, human gaze also includes noise and redundancies that may transfer poorly to VLMs. In this work, we investigate how human gaze data can be effectively leveraged to improve visual reasoning in VLMs, while maintaining alignment with human behavior. We collect 7,202 human gaze trials from a diverse and challenging set of visual reasoning tasks, and train grounded reasoning models with different forms of gaze guidance. Across evaluations in chart understanding, visual search, and counting, we find that selective use of human gaze consistently outperforms supervision that incorporates the full gaze sequence, achieving greater human alignment while maintaining performance competitive with leading grounded-reasoning baselines. These gains persist through reinforcement learning, while gaze-guided models remain more spatially aligned with human behavior. Behavioral analyses further show that gaze supervision induces persistent differences in exploration and verification. Our results suggest that the value of human gaze lies in transferring its exploration structure without requiring complete imitation of human behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.