acceptodds
Under review as a conference paper at ICLR 2027

SENRI: Understanding the Space-Event Nexus in Real-Life Videos

Abstract

Human memories are organized in a way that connects events with their corresponding spatial contexts. An AI assistant should also be able to memorize and recall the nexus of events and spaces. Thus we introduce **SENRI**, a task that studies this space-event connection in long, continuous real-life egocentric videos. It contains six task types in two directions: *EventSpace*, which recovers spatial knowledge from event cues, and *SpaceEvent*, which retrieves events from regions and trajectories. To evaluate this, we build **SENRI-Bench**, containing 500 _manually annotated_ multiple-choice questions over 10 uninterrupted _self-recorded_ videos with around 25.3 hours of content. Also, we evaluate three typical long video understanding paradigms: direct perception, evidence seeking, and persistent memory. The task remains challenging across all three paradigms, highlighting substantial room for improvement. Then we show two representative failure modes of existing memory methods and address these evidence gaps with **SENRI-Agent**, a controller that uses _Target-Guided Visual Check (TGVC)_ to propose local visual check tasks for parallel verifiers and integrate the reports into a working memory to help with reasoning. Built upon existing memory systems, SENRI-Agent improves the performance of existing baselines without modifying their stored memories. We hope SENRI will provide valuable insights into how real-life egocentric understanding should be evaluated, reveal the gaps in existing approaches, and offer a possible direction for improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.