acceptodds
Under review as a conference paper at ICLR 2027

StreamVLM: Adapting Offline Vision-Language Models for Streaming Understanding

Abstract

Recent vision-language models (VLMs) have demonstrated strong capabilities in video understanding. However, most existing VLMs are designed for offline inference, where the entire video sequence is jointly processed during reasoning. Directly adapting such offline VLMs to streaming tasks is challenging, since streaming scenarios require the model to incrementally process incoming frames without access to future visual content. In this work, we propose StreamVLM, a training-free framework for adapting offline VLMs to streaming video understanding. StreamVLM introduces Event-based Streaming Encoding, which dynamically constructs KV-cache memories according to motion-aware temporal boundaries to preserve coherent local context during streaming inference. In addition, we propose a KNN-based KV-Cache Retrieval mechanism that models neighborhood relationships among historical memory entries to enable context-aware retrieval over correlated KV representations. Extensive experiments on both offline and streaming VideoQA benchmarks demonstrate that StreamVLM consistently improves video reasoning performance over existing streaming VLM baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.