When to Trust the Eyes: Learning Adaptive Gaze Reliance in Egocentric Video
Abstract
Gaze recorded by wearable devices reveals where the wearer attended, and prior work exploits it for egocentric video understanding under the implicit assumption that it is useful regardless of the task. However, we show that not all gaze is useful for question answering, the core function of a wearable assistant. Rather, its usefulness depends on the question and the scene context, so the same fixation can be useful for one question yet noise for another. Accordingly, how much to rely on gaze should be decided for each question and frame, trusting gaze where it is useful and suppressing it where it acts as noise. To this end, we propose **Re**liance on **Gaze** (ReGaze), a method that learns this reliance. Specifically, we design the Adaptive Gaze Reliance Estimation (AGRE) module, a lightweight network trained without any label of gaze usefulness, using a ranking loss that favors reliance on gaze where it helps localize question-relevant frames and a noise suppression loss that penalizes it where it hurts. To train the AGRE module, we further build a Gaze-aware QA Curation (GQC) pipeline that generates questions paired with their relevant intervals from gaze-recorded videos. On three egocentric QA benchmarks, ReGaze outperforms state-of-the-art frame samplers and naive uses of gaze. Beyond performance, ReGaze captures the underlying behavior of gaze without being trained on it: its reliance decreases as the gaze shifts rapidly or lies in the periphery. It also excels on WearerQA, our evaluation set of questions written by the wearers themselves.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.