acceptodds
Under review as a conference paper at ICLR 2027

StreamBank: Unified Generation–Memory Inference for Efficient Streaming Video Understanding

Abstract

Streaming video understanding requires Video Large Language Models (Video-LLMs) to generate timely responses while preserving relevant historical information from continuous video streams. However, autoregressive response generation incurs substantial decoding latency, while continuously growing visual–textual histories introduce increasing computation and memory costs. Even with speculative decoding, candidate generation often remains sequential, while existing memory policies typically rely on visual similarity or recency rather than prediction-relevant evidence to retain historical events. To address these limitations, we propose StreamBank, a unified inference framework that uses target-model verification to accelerate response generation and guide historical event retention. StreamBank comprises two complementary components. First,Reuse-Augmented MTP Speculative Decoding predicts multiple future tokens in parallel and reuses previously verified responses as candidates when temporal continuity permits, reducing drafting overhead while retaining target-model verification. Second, Speculation-Guided Event Memory repurposes verifier-gain and correction signals from the same verification process to guide event promotion, retention, and eviction within a fixed-capacity memory bank. Together, these components connect efficient generation with prediction-aware memory management through shared verification feedback. Experiments on ProactiveVideoQA, OVO-Bench, and StreamingBench show that StreamBank largely preserves streaming video understanding performance under substantial context compression while achieving up to a decoding-throughput speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.