TRACE: Novelty-Aware Spatiotemporal Evidence Retention for Streaming Video Understanding
Abstract
Streaming video understanding requires vision-language models to process evolving streams causally and preserve visual evidence for queries arriving at arbitrary times. However, models must make irreversible retention decisions under a bounded memory budget, without access to future observations or queries. Existing methods often treat temporal and spatial redundancy separately or rely on attention-based importance scores, potentially discarding useful evidence before complementary cues are jointly considered. To address this challenge, we propose TRACE, a training-free framework for novelty-aware spatiotemporal evidence retention. Following an observe–score–retain procedure, TRACE maintains a causal hierarchical visual memory that preserves recent observations at full granularity while characterizing visual tokens through two complementary query-agnostic signals: temporal novelty and spatial distinctiveness. Their joint evaluation enables adaptive evidence retention without relying on attention weights, while temporal gap-aware positional re-encoding supports long-horizon streaming inference. Extensive experiments across multiple VLM backbones demonstrate that TRACE achieves state-of-the-art performance among the compared open-source methods on streaming video understanding benchmarks, while preserving 91.1% of the corresponding backbone performance with only 8.2% of the visual tokens on long-video benchmarks. TRACE further maintains bounded visual KV memory and stable query latency over hour-scale streams.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.