EgoPass-Bench: Identifying Objects That May Affect the Wearer's Passage in Egocentric Video
Abstract
Understanding an egocentric environment requires reasoning about how surrounding entities affect the observer's actions. Existing evaluations of multimodal foundation models (MFMs) examine observer localization, motion trajectories, and action feasibility, but leave the discovery and interpretation of evolving observer–entity interactions insufficiently assessed. We introduce EgoPass-Bench, a benchmark for observer-relative interaction understanding in real-world egocentric video. EgoPass-Bench comprises 95 manually reviewed egocentric clips recorded using Ray-Ban Meta (Gen 2) smart glasses and 577 distinct questions across five tasks assessing motion perception, interaction judgments about specified entities, and relevant-entity discovery without target cues. Our evaluation of eight MFMs reveals a pronounced discrepancy between motion perception and interaction understanding. On EgoPass-Bench, GPT-6 Astra leads on most evaluation metrics and achieves approximately 90% accuracy in recognizing both observer and external-object motion, yet it achieves only 60.00% on relevant-entity discovery without target cues. Further analysis reveals frequent inclusion of entities that neither pose a plausible physical conflict nor constrain passage, showing that accurate motion recognition does not ensure reliable relevant-entity discovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.