Fusing Complementary Multi-view Feature for Screen-based Eye Tracking
Abstract
Current multi-view gaze estimation remains limited by existing datasets, insufficient exploitation of complementary cross-view information, and evaluation focused primarily on average gaze error. We address these limitations through a more systematic study of multi-view gaze estimation. First, we introduce PrismGaze, a new dataset with over three million images, capturing continuous head-pose variation for the same gaze targets. Second, we propose PrismFusion, a multi-view feature fusion framework based on region partitioning, which masks complementary image regions across views during training to encourage effective cross-view information integration. Third, we develop a broader evaluation framework that examines the effects of camera number and placement, target location, viewing depth, and unseen viewpoints. Our experiments show that the primary benefit of multi-view gaze estimation comes from compensating for poorly observed views with cameras providing more favorable viewpoints. PrismFusion remains robust to changes in viewing depth and unseen viewpoints, while achieving strong calibration-free performance on seen views. Together, our dataset, method, and evaluation provide a more comprehensive foundation for studying multi-view gaze estimation in realistic settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.