ConcurTree: A Novel Memory Structure for Long-Video Streaming
Abstract
Vision-Language Models (VLMs) have advanced rapidly in visual perception and reasoning, yet long-video understanding in streaming applications remains challenging due to strict memory constraints and incomplete reasoning over extended temporal horizons. Existing streaming VLMs face a key trade-off: structured query-agnostic systems track objects or locations with unbounded token growth, whereas fixed-budget architectures sequentially merge adjacent events, losing interleaved activity threads when tasks are interrupted and resumed. To address these limitations, we introduce ConcurTree, a novel query-agnostic memory framework that maintains distinct trees under a fixed token budget. ConcurTree's memory structure organizes events into activity threads causally with an exception-based routing mechanism, while handling memory constraints via similarity-weighted token merging that compresses historical evidence in-place rather than evicting or flattening it. ConcurTree achieves state-of-the-art long-horizon performance among peer streaming systems across three held-out benchmarks under controlled evaluation. On our primary pooled long-horizon endpoint (n=1,054), pooled across three benchmarks, ConcurTree reaches 50.4% accuracy, compared with 43.1% for the strongest matched fixed-budget streaming baseline. On EgoLifeQA and Ego-R1, ConcurTree reaches 48.6% and 39.5% when evidence is more than 24 hours old, respectively, and achieves 52.3% on TeleEgo's long-horizon categories. These results highlight the importance of addressing activity structure for long-horizon reasoning in bounded streaming video.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.