LookAnything: A Generalist Foundation Model for Gaze Target Detection
Abstract
Human gaze provides a rich window into attention and intent, making its understanding essential for applications ranging from social signal analysis to human-robot interaction. Existing methods typically employ highly customized architectures relying on auxiliary inputs like head crops or depth maps, causing cascading errors and inference overhead. Trained on single domains and confined to isolated gaze target detection with heatmap outputs, they lack flexibility and generalization for diverse interactive scenarios. To bridge this gap, we present LookAnything, a generalist foundation model unifying diverse gaze-centric tasks. Benefiting from large-scale multi-task cross-domain training, LookAnything flexibly adapts to various settings, generating diverse outputs (points, boxes, heatmaps, and language) and supporting broad gaze-centric localization, description, and reasoning. At its core, we introduce learnable gaze-prior queries to promote gaze-aware representation learning, eliminating the need for auxiliary external inputs. To facilitate this paradigm, we construct HoliGaze, a comprehensive corpus and benchmark comprising 311K question-answer pairs across 5 capability dimensions and 9 distinct tasks. Experiments demonstrate that LookAnything consistently outperforms 33 leading vision-language models on holistic gaze understanding, while remaining highly competitive with dedicated specialists across 5 conventional target detection benchmarks. Deployment on a Unitree G1 robot further validates its strong out-of-domain generalization and flexible adaptation in real-world scenarios. Our codes, model, and dataset will be released. The project website is available at https://anonymous1541paper.github.io/iclr2027-1541.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.