acceptodds
Under review as a conference paper at ICLR 2027

Event-ReHead: Event-Aware Head Retrieval and Memory-Augmented Streaming Video Understanding

Abstract

Vision Language Models (VLMs) have achieved strong performance in video understanding, while streaming video question answering remains challenging due to the need for accurate retrieval under strict memory and latency constraints. Existing KV-cache-based methods mainly organize video memory as independent visual tokens, which ignores underlying temporal semantics of videos and leads to fragmented retrieval, failing to capture continuous evidence . We propose Event-ReHead, a training-free framework that improves streaming video understanding by introducing event semantics into KV-cache retrieval. Our method preserves the efficiency of frame-level retrieval while reorganizing retrieved evidence into an *event skeleton*, a compact and coherent representation capturing the underlying temporal structure. Building on this, we introduce a query-specific attention head selection strategy to construct more discriminative retrieval representations, enabling precise alignment with question-relevant visual signals. We further incorporate a lightweight interaction memory that selectively reuses prior question-answer context, providing useful temporal priors without introducing noise. We evaluate Event-ReHead on four streaming and three offline benchmarks, demonstrating consistent improvements across diverse VLM backbones. In streaming settings, it achieves up to +4.1% gains, while also yielding consistent improvements in offline scenarios, with an average gain of around +2-3%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.