HippoStream: Hippocampal Memory for Streaming Video Understanding
Abstract
The hippocampus is associated with the rapid encoding of recent experience into a compact representation that other regions can later draw on. Streaming Video LLMs face the same problem with no comparable mechanism: they either keep a sliding window of recent frames and forget the past, or grow a key–value cache that scales with stream length and breaks on long videos. HippoStream is a small parametric token bank that the inference loop rewrites with functional gradient steps per chunk while the Video LLM trunk stays frozen. The bank plays the role of hippocampus, the frozen trunk the role of neocortex. Its initial state, writer head, and read projection are meta-learned offline; per-chunk cost is independent of stream length, and the bank's footprint stays constant across videos of any length. On OVO-Bench with Qwen3-VL-8B, HippoStream improves the Backward-summary cluster (Action-State Inference, Hold-out reasoning) by over the recent-window baseline. The gain traces to a single design choice in the WRITE step: anchoring the bank's output positions to each chunk's first vision token preserves the event-specific evidence that backward retrieval needs. Combined with offline meta-learning of the slow weights, this gives a streaming Video LLM reliable access to past events long after the recent window has scrolled by. We release code, meta-trained checkpoints, and the training-data replay buffer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.