GazeSieve: Gaze-Guided Keyframe Selection for Egocentric Video Question Answering
Abstract
Egocentric video question answering often depends on a few evidence moments: an attended object, a brief state change, or an earlier observation that explains the query. The bottleneck is therefore selecting the few frames that contain sufficient evidence under a small visual budget, yet downstream answer accuracy alone cannot reveal whether selection succeeded. Recent video-language agents can perform query-conditioned search, but remain gaze-free and may miss moments humans naturally attend to. We present GazeKeyBench, a direct keyframe-selection benchmark that annotates acceptable evidence windows, required evidence groups, and alternative sufficient evidence sets, judging whether selected frames cover the evidence rather than land near one timestamp. We also present GazeSieve, a training-free gaze-guided selector that turns noisy gaze into fixation proposals, extracts query-relevant concepts, searches high-resolution gaze-centered crops under a gaze-informed temporal prior, and returns a relation-aware evidence set to a frozen MLLM. Under a unified controlled protocol that freezes the answerer, decoding, prompt, rendering, and frame budget, replacing uniform sampling with GazeSieve raises StreamGaze Macro-8 from 58.0 to 60.8 (+2.8, p < 0.01) on a frozen 4,700-unit roster. On GazeKeyBench's 2,607 human-annotated cases from 99 videos, GazeSieve reaches 56.0% Evidence Recall@4, outperforming a strong gaze-free Agent(VLM) baseline at 52.0%. A paired timing control further shows that preserving measured gaze timestamps adds 1.9 points (p = 0.0247). These results suggest that measured gaze provides complementary evidence-selection signal for egocentric video QA beyond query-conditioned video search alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.