acceptodds
Under review as a conference paper at ICLR 2027

EgoTrace: Coupling Semantic Proximity and Temporal Continuity for Egocentric Video Question Answering

Abstract

Long-horizon video question answering often begins with an incomplete cue: a question names part of an experience without specifying the evidence needed to answer it. Retrieving isolated records by similarity is fragile here, because similar encounters recur and the record that matches a question may not be the one that answers it. Inspired by contextual reinstatement in episodic memory, we introduce EgoTrace, which keeps original captions and aligned speech unchanged and links them to a scene trace: source-linked descriptions of how the surrounding situation develops. A Controller iteratively chooses among searching for records, reinstating the context around a retrieved record, and opening original evidence; a separate Reader answers from the opened sources and their linked context. Because the trace guides navigation rather than defining retrieval units, the extent of recovered context adapts to each question instead of being fixed by a window or pre-segmented events, and opened evidence can turn an incomplete match into a more precise cue. The trace is built in a single chronological pass and can be extended as footage arrives. EgoTrace achieves the highest reported scores of 74.6% on EgoLifeQA and 33.4% on MM-Lifelong Month validation, exceeding prior results by 3.4 and 8.9 percentage points, respectively. Ablations show that trace descriptions add value beyond temporal access alone, while direct source retrieval outperforms tested graph and semantic-memory additions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.