ProactEgoMem: Benchmarking Proactive Memory in Streaming Egocentric Video
Abstract
Today's multimodal assistants are typically prompted to interpret the visual evidence presented to them. To offload the cognitive burden of remembering across long horizons, wearable assistants should proactively acquire and maintain useful memory as the user's surroundings evolve. Remembering the ambience on the user's behalf requires identifying what matters, understanding the scene in real time, and initiating memory maintenance in the background while the user focuses on other activities. We introduce ProactEgoMem, a benchmark that evaluates proactive memory in streaming egocentric video. Three tasks ground the evaluation in ambient environmental needs: Information Capture (IC) preserves transient yet dense visual information, State Tracking (ST) follows changing object states, and Location Tracking (LT) records object movements. The evaluation protocol spans two levels of proactiveness: the Notification Level asks for one timely notification that satisfies the user's intent; the deeper Intervention Level asks for memory maintenance through an active start-to-stop tool lifecycle on the user's behalf. Two parallel tracks, ProactEgoMem-Daily and ProactEgoMem-Sci, contain 300 and 100 clips for daily and scientific use cases, respectively, paired with 1000 and 500 questions in ProactEgoMem-QA, which measures the quality of proactively acquired memory. Across proprietary, open-source, video-language, and streaming baselines, even the strongest models fail to capture memory proactively with accurate timing and content. ProactEgoMem provides a rigorous evaluation foundation for multimodal proactive assistance and co-existing video agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.