TRACE: Temporal Resource-Aware Chains of Evidence for Long-Range Video Understanding
Abstract
Streaming long-video understanding requires a model to answer user queries at any moment while maintaining visual and conversational context across multiple rounds. However, existing methods either reduce memory through flat representations that lose the temporal relations among events, or maintain such relations in structures that grow continuously with the stream. To address these issues, we propose TRACE, a training-free online framework that organizes streaming memory into queryable chains of evidence with resource-aware retention and retrieval. During ingestion, we compare each observation with its predicted semantic state and retain surprising event changes in full detail while merging repeated content. The retained observations are then parsed into subject-action-object event nodes and organized into a spatio-temporal chain that preserves entity trajectories and event relations. When a query arrives, an intention-driven resource-aware router formulates evidence access as a budgeted selection problem and dynamically allocates the limited resource budget to maximize the utility of complementary evidence for response generation. Extensive experiments show that our framework achieves superior performance with accuracy of 76.2% and answer-quality score of 3.88 on StreamBench. It also demonstrates high throughput of 94.05 FPS with low latency of 3.36 s and maintains generalization capabilities on the other complementary benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.