GazeGGT: Geometry-Grounded Transformer for Uncalibrated Multi-View Gaze Estimation
Abstract
Robust 3D gaze estimation remains challenging because monocular appearance cues are inherently under-constrained: eye visibility, head pose, and self-occlusion can make the same gaze direction appear substantially different across views. Multi-view observations offer a natural remedy, but existing multi-view gaze estimators typically depend on fixed camera layouts, calibrated cameras, or test-time adaptation, limiting their applicability in flexible real-world environments. We study the underexplored problem of uncalibrated multi-view gaze estimation, where multiple synchronized face observations are available but camera intrinsics, extrinsics, and view configurations are unknown at inference. We propose GazeGGT, a geometry-grounded transformer for unconstrained, calibration-free multi-view gaze estimation. GazeGGT encodes an arbitrary view number of face crops into per-view features and employs alternating multi-view attention to fuse information across views, followed by separate decoders for per-view 3D gaze directions and camera poses. Unlike prior multi-view methods, GazeGGT requires no camera parameters and naturally supports a variable number of input views, enabling flexible camera configurations. We further introduce a Geometry-Grounded Supervision Paradigm that jointly supervises gaze and camera pose predictions, providing the geometric grounding necessary for accurate calibration-free gaze estimation. Experiments show that GazeGGT achieves competitive performance in both single-view and dual-view settings and consistently benefits from increasing numbers of input views. Ablation studies further demonstrate the importance of the Geometry-Grounded Supervision Paradigm for calibration-free estimation. Our results establish a practical and extensible approach to gaze estimation in unconstrained multi-camera environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.