EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes
Abstract
Multimodal large language models (MLLMs) deployed in devices such as smart glasses need to identify appropriate next actions from the user's perspective. However, some questions in existing evaluations based on first-person visual inputs can be answered based solely on the camera wearer's body cues. When other people's actions are absent from the answer choices, a correct response provides limited evidence that the model distinguishes the user's situation from those of others, so it is difficult to assess action selection from the user's perspective separately from input understanding. To address this limitation, we use images of multiple people that contain no body cues from the camera wearer and ask models to respond as if they were directly viewing the scene. We define Egocentric Action Selection (EAS) as the task of selecting an appropriate action from the user's perspective in a given situation and introduce EgoGapBench, a diagnostic benchmark for EAS with 1,000 multiple-choice questions on 440 COCO images. EgoGapBench includes other people's actions as distractors when duplicating those actions would be inappropriate in the given situation. All evaluated open- and closed-source MLLMs achieve substantially lower accuracy than the human baseline, and for most models errors concentrate on the actions of other people in the scene. Moreover, after instruction tuning on first-person visual input data, all five open-source models score nearly the same on an existing benchmark of first-person visual inputs ( relative) but drop on EgoGapBench ( relative).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.