Understanding What You Point At: Spatial Guidance for Reference Disambiguation in Egocentric Video
Abstract
Pointing gestures link linguistic references to physical objects, providing essential cues for egocentric AI assistants to interpret user intent and interact naturally. However, current multimodal large language models (MLLMs) struggle to distinguish intended targets from distractors in spatially complex and dynamic scenes. Existing benchmarks provide limited systematic coverage of these ambiguities. We introduce PointConfusion, to our knowledge the first egocentric video understanding benchmark designed around a taxonomy of spatial and dynamic pointing ambiguities. Its 5,000 synthetic videos enable systematic evaluation of five ambiguity categories through a unified target-selection task. Our method, PointMod, learns spatial reference heatmaps from hand geometry and uses them to modulate visual features, explicitly incorporating pointing information into the model’s visual representation. Extensive evaluations on PointConfusion demonstrate that PointMod outperforms a range of advanced MLLMs and achieves state-of-the-art performance. PointMod-8B reaches 88.33% target-selection accuracy, improving upon task-fine-tuned InternVL3.5-8B by 8.13 percentage points. The dataset and code will be publicly released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.