The Right Memory at the Right Time: Memory-Driven Proactive Assistance
Abstract
As assistants move from reactive answering toward proactive intervention, memory changes role: it no longer only supplies the answer to a query, it determines whether and when to speak at all. Prior work in this setting hands the system the relevant fact in context; we study the regime that makes the problem hard, where the evidence justifying an intervention must be mined from a long-horizon personal memory. We introduce the memory-driven proactive assistance (MPA) benchmark, pairing egocentric video with year-long, distractor-rich user memories ( 618K tokens on average), human-verified trigger windows, and a protocol that scores content and timing jointly. Every model we evaluate scores near zero, including frontier models given summarized or retrieval-augmented memory. A controlled study explains why: with oracle evidence delivered inside the trigger window a frontier model reaches 96%; letting it choose the timing drops it to 38%, and replacing oracle evidence with retrieval drops it to 15%. Content and timing are separately necessary, and missing either is catastrophic. Guided by this, we build Gated Proactive Memory (GPM), a plug-and-play harness that surfaces memory from a graph-organized store and gates it down to what is worth interrupting for, plus a lightweight RL stage that curbs early and repeated firing. GPM lifts strict accuracy from near zero to 37.0% on Qwen3.8-27B and, with no training at all, 38% on GPT-5.3-Chat; RL post-training reaches 42.8%. We will release the MPA benchmark upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.