Beyond Saliency: Token Memory via Marginal Contribution for Streaming Video Understanding
Abstract
Streaming video understanding requires processing an unbounded video stream within bounded memory while answering queries whose timing and content are unknown in advance. Because evicted tokens cannot be recovered, memory management should preserve information not already covered by the current memory rather than simply retain individually salient tokens. Existing saliency-based methods do not explicitly model such redundancy, while age-stratified methods constrain competition across temporal tiers. We propose SMART, a training-free, query-independent streaming token memory that retains tokens according to their marginal contribution beyond what the current memory already covers. SMART groups consecutive frames into content-based events to distinguish within-event redundancy from observations across events. Within each event, cross-frame nearest-match weights estimate how much a token is already covered by other retained frames, and modulate an intra-frame determinantal point process (DPP) that measures token complementarity. Because marginal contribution changes as memory evolves, SMART rescores surviving tokens whenever a new frame arrives. An age-augmented DPP kernel further combines contribution with recency for unprotected tokens, while preserving the most recent frame for real-time perception. With Qwen3-VL-8B, SMART achieves 68.6% OVO-Avg and 82.4% accuracy on StreamingBench, the highest aggregate results among the streaming methods evaluated.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.