ResKV-Tree: Residual KV-Caching Tree For Streaming Video Understanding
Abstract
Streaming video understanding requires storing historical evidence for questions that are unknown when frames arrive, making it essential to reduce memory redundancy while preserving potentially useful details. Existing methods rely on granularity-agnostic storage that compresses all evidence uniformly, making it difficult to preserve information at different granularities. To address this, we propose Residual KV-Caching Tree (ResKV-Tree), a hierarchical KV-caching structure for effective and efficient evidence retrieval. During the construction of ResKV-Tree, we select nodes to merge based on semantic similarity, temporal proximity, and evidence saliency, then apply KV non-maximum suppression (KV-NMS) to remove redundant KV pairs using key-based saliency estimation and value-based redundancy suppression. To form the hierarchical structure, we recursively promote salient KV pairs to parent nodes and keep residual ones in children. During question answering, best-first search is employed to explore hierarchical paths starting from the roots and efficiently retrieve relevant evidence at different granularities. Extensive evaluations on two streaming and three offline video understanding benchmarks demonstrate the superiority of ResKV-Tree over baselines and recent methods across three different backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.