Look Less, Remember More: Training-free Dynamic Encoding with Hierarchical Visual Memory for Streaming Video Understanding
Abstract
Streaming video understanding must answer questions about an evolving video stream within restricted compute budgets. However, the visual tokens grow linearly with video length, and most of them repeat the content already seen. Existing studies typically prune visual tokens in the large language model under a query-aware criterion, missing the opportunity to remove redundancy at the early encoding stage. To address this issue, we propose , a novel training-free and query-agnostic framework that relies solely on the intrinsic cross-modal semantics of a pretrained Video-LLM, which can dynamically encode the video frame, removing spatiotemporal redundancy, and effectively manage and utilize historical visual information. First, we propose a dynamic encoding approach via inter-frame temporal novelty and salience-based spatial token pruning to improve the efficiency of video representation and eliminate spatiotemporal redundancy. Second, to effectively organize and leverage historical visual information, a diversity-preserving hierarchical memory mechanism is incorporated, which can retain the most informative visual content in the memory bank without the requirement for query guidance. Extensive experiments across online and offline long-video benchmarks demonstrate that achieves superior video understanding performance with limited memory usage and lower latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.