acceptodds
Under review as a conference paper at ICLR 2027

Learning to Gaze Beyond the Face: Unsupervised Transfer from Cropped-Face Experts to Full-Frame Driving Scenes

Abstract

Large-scale gaze estimation datasets are predominantly constructed from cropped face images, enabling accurate gaze prediction but limiting direct deployment in real-world driving scenarios, where cameras capture full-frame images containing complex backgrounds and contextual information. This discrepancy creates a fundamental challenge: how can gaze knowledge learned from cropped faces be transferred to full-frame driving scenes when target-domain gaze annotations are unavailable? In this work, we present an unsupervised framework for cropped-face-to-full-frame gaze transfer. Rather than forcing full-frame representations to directly mimic cropped-face features, which can suppress useful contextual information and cause degenerate predictions, we retain a pretrained cropped-face gaze expert and use it to provide privileged gaze knowledge during training. A full-frame model is subsequently optimized with self-supervised constraints that encourage gaze-relevant representation learning while preserving the intrinsic variation of gaze predictions. In particular, an anti-collapse constraint is introduced to maintain the batch-level gaze distribution and prevent the model from converging to nearly constant predictions. This design allows the full-frame model to learn gaze-sensitive representations without requiring target-domain gaze labels or explicit feature-level alignment between cropped faces and full frames. Experiments on in-vehicle gaze estimation demonstrate that the proposed approach substantially improves over direct application of the cropped-face expert and achieves stable gaze estimation from raw full-frame inputs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.