Remembering by Place: Persistent Visual Memory for Streaming Video Agents
Abstract
Streaming video models promise always-on assistants that can later recall what they have seen, yet their memory is bounded by context: observations evicted from the KV cache are discarded rather than preserved as a queryable record. Existing memory agents for long video typically fill this gap by captioning incoming observations and storing text. We argue that this caption-first design is mismatched to visual-spatial experience: it commits visual evidence to language before the question is known, scales writer-LLM call with the observation stream, and does not link revisits to the same place across time. We introduce PlaceMem, a query-blind, training-free persistent memory agent that serves as the backend of a streaming video model. Taking the surprise-selected observations of Cambrian-S as its front end, PlaceMem organizes them into persistent place records using visual place and semantic features, and derives language from accumulated place evidence rather than from each observation. A tool-using reader retrieves across both the textual and visual records. On VSI-Super Recall, PlaceMem improves average accuracy on 10–240-minute streams by 31.7 points over Cambrian-S alone, even with a lightweight 4B reader, and writing language per place rather than per observation cuts writer calls by 38% at competitive accuracy. On VSI-Super Count, it improves Cambrian-S's MRA from 40.7 to 48.2. These results suggest that caption-first memory need not be the default for visual-spatial memory: captions remain strong for semantic recall, but visual features and place organization supply grounding and cross-observation association that text alone does not encode.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.