TimeGraph: Temporal Graph Memory for Long-Horizon Video Question Answering
Abstract
Long-horizon video question answering requires reasoning over information distributed across distant observations, where entities may persist while their states and relations change over time. Retrieving individual frames does not explicitly preserve these temporal dependencies, leaving entity continuity, temporal extent, and state changes to be reconstructed during answering. We propose TimeGraph, a temporal graph memory that organizes frame-level scene-graph observations into persistent entities, temporally bounded state and relation episodes, and explicit transition boundaries on a shared time axis. Given a question, TimeGraph uses entity and temporal constraints to retrieve a relevant temporal subgraph and convert its structure into answer-relevant evidence. We evaluate TimeGraph on NAVQA, NExT-QA, and OpenEQA, where it improves over previously reported results by 6.2, 3.1, and 1.9 points, respectively. Compared with frame-local lexical retrieval, temporal graph memory further improves performance by 35.25, 13.39, and 21.14 points across the three benchmarks. Further analysis shows that organizing visual evidence according to question demand can provide additional gains, although the effect varies across datasets and question types. These results demonstrate the value of establishing explicit temporal structure before question-time retrieval for long-horizon VideoQA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.