acceptodds
Under review as a conference paper at ICLR 2027

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

Abstract

Streaming video understanding is essential for real-time embodied and assistive systems, but processing open-ended streams with multimodal large language models (MLLMs) is computationally expensive. Existing methods reduce visual tokens or stored context, yet still apply full-depth prefill to admitted frames, causing repeated computation and depth-proportional KV-cache growth. We observe that shallow MLLM layers already provide effective signals for retrieving question-relevant evidence. Building on this observation, we propose **ShallowStream**, which decouples lightweight stream processing from selective full-depth answering. During streaming, ShallowStream processes frames only through shallow layers, using their KVs as a historical index with optional cluster compression. At query time, a text-only gate activates cross-layer token voting and diversity-aware retrieval when needed, after which only the selected history and recent context undergo full-depth prefill. ShallowStream reduces per-frame prefill and 20-second end-to-end latency by up to **52.1×** and **15.3×**, respectively, while improving average scores by approximately **1.7 points** over strong existing baselines. Our code is available at [https://anonymous.4open.science/r/ShallowStream/](https://anonymous.4open.science/r/ShallowStream/).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.