StreamIndex: Decoupling Indexing from Encoding for Scalable Streaming Video Understanding
Abstract
Serving concurrent video streams under a fixed GPU budget requires sustained ingestion and timely responses to unpredictable queries. Existing streaming video MLLMs typically encode incoming frames before deciding which information to retain or retrieve, making continuous visual computation a scalability bottleneck. We propose StreamIndex, a training-free framework that decouples stream-time indexing from query-time visual encoding. During ingestion, lightweight CPU-side statistics group similar frames into clusters, and only one representative per finalized cluster is encoded for retrieval. When a query arrives, the system retrieves relevant clusters and recovers details through motion-based frame selection and selective reuse of stable spatial features. This design reduces continuous encoding while controlling the additional work on the query path. Experiments on the Qwen-VL and LLaVA-OneVision families show improved multi-stream scalability with competitive accuracy on OVO-Bench and StreamingBench. In a 16-stream 1080p query-load test at two queries per minute per stream, StreamIndex achieves a P95 response latency of 5.57 seconds, while existing state-of-the-art baselines either violate the real-time bound or run out of memory. The results expose a practical trade-off between ingestion cost, query-time computation, and answer quality. The code is available at https://anonymous.4open.science/r/StreamIndex-7D6E.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.