WhenMemBench: Benchmarking Implicit Memory Need Recognition During Task Execution
Abstract
Long-term memory helps only when an agent recognizes that hidden history is needed, accesses it when its relevance becomes observable, and applies the right evidence. Existing memory benchmarks often reveal historical context, precompute retrieval results, or issue explicit recall requests; therefore, they evaluate memory use after the access decision has already been made. However, real users usually provide underspecified requests and may not know which prior detail matters. We formalize this gap as Implicit Memory Need Recognition: without an explicit memory recall request from the user, the visible state alone becomes insufficient for the agent to fulfill a task-critical requirement that implicitly calls for evidence from the hidden interaction prefix. We introduce WhenMemBench, an executable benchmark for this problem. To establish the benchmark, we convert multi-turn tool trajectories into auditable tasks through automated trajectory mining and naturalization followed by human-in-the-loop validation. The benchmark consists of 164 tasks, distinguishing Query-Triggered tasks, where historical relevance is inferable from the initial request, and Execution-Triggered tasks, where it emerges after workspace or tool observations. We further propose Step-Level Proactive Memory, a baseline strategy that uses a separate memory agent to selectively deliver relevant history at each execution step, and evaluate it among seven memory strategies on WhenMemBench using GPT-5.6. Our strategy achieves 43.29% overall task success, compared with 27.44% for an agent-controlled memory tool and 18.29% for Always-On RAG. These results reveal that implicit memory need recognition remains challenging for agents, even when the relevant history is retrievable from external memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.