acceptodds
Under review as a conference paper at ICLR 2027

EgoSpect: Egocentric Spatial Perception Event-Cascade Testbench

Abstract

EgoSPECT is a first-person audio-visual benchmark for evaluating event-driven embodied intelligence across perception, priority assessment, and behavioral decision-making. Unlike benchmarks that evaluate these capabilities in isolation, EgoSPECT uses explicitly grounded question chains to measure the full event cascade while localizing failures to individual stages. We collect approximately 60 hours of egocentric recordings across six indoor and outdoor scenarios and evaluate three multimodal large language models under audiovisual, audio-only, video-only, and text-only conditions. Across the evaluated tasks, perception is the primary bottleneck: models frequently miss weak or masked events, confuse event categories, and over-report ambient sounds as hazards. In the complex multi-event scenario, end-to-end accuracy of 27.1%–62.8% rises to 69.0%–79.2% under conditional evaluation, while the three tested models reach 100% on 96 closed-form downstream questions when the ground-truth event pair is supplied as a premise. Eight-way spatial direction estimation remains near chance, but controlled DOA intervention substantially improves performance, with Oracle DOA revealing further headroom. These results show that robust embodied perception requires not only recognizing events, but also distinguishing relevant events from ambient sound and jointly grounding event semantics with spatial location under noisy and concurrent conditions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.