acceptodds
Under review as a conference paper at ICLR 2027

Learning World-Space Representations of Human Attention from Uncalibrated Views

Abstract

Human gaze target estimation (GTE) remains challenging due to the limited camera field of view. While multi-view setups can mitigate this issue, prior attempts either operate in fixed camera setups or heavily rely on explicit bounding boxes. In this work, we present GazeWorld, a unified framework dedicated to learning world-space representations of human attention directly from uncalibrated multi-view images in the wild. By constructing a unified 3D-aware representation that jointly encodes scene geometry and semantics, GazeWorld leverages pre-trained geometry priors, establishes prompt-based subject association, and predicts per-view gaze targets using features expressed in a shared, estimated 3D coordinate frame. This world-space modeling allows our approach to process a variable number of uncalibrated views, helping resolve ambiguities across views. To support this paradigm, we curated a large-scale dataset from Ego-Exo4D containing over 260K verified multi-view GTE groups, alongside a self-collected in-the-wild benchmark featuring challenging head-invisible scenarios. Experiments show strong GTE performance on multiple benchmarks under both single-view and multi-view scenarios, and demonstrate consistent gains with increasing input views. Code, models, and benchmarks will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.