From Memory to Action: Benchmarking Long-Term Agent Memory in Personal Assistant Environments
Abstract
Always-on personal assistants are an emerging application of agents powered by large language models (LLMs). Unlike conventional assistants, they operate over long-term, complex streams of user activities and agent interactions, accumulating information that is essential for tasks with ambiguous instructions, evolving user preferences, prior feedback, and reusable experience. This places substantial demands on agent memory, not only to form and maintain useful memories, but also to correctly apply them during downstream task execution. Existing benchmarks, however, either evaluate memory largely in isolation from end-task success or rely on simplified interaction histories that do not reflect the complexity of sustained daily use. As a result, there remains a gap in evaluating agent memory end to end, from memory formation to its eventual effect on task completion. To address this gap, we introduce a comprehensive memory benchmark comprising 100 tasks with fine-grained target-memory attribution, enabling separate diagnosis of Formation, Availability, Application, and Outcome. Evaluations across different LLMs show that the largest performance loss occurs after relevant memory becomes available but before it is translated into the intended action, indicating that memory utilization, rather than storage or retrieval alone, is a major bottleneck. Controlled memory-swap experiments further reveal that the utility of a generated memory depends strongly on the downstream executor that uses it. We hope these insights can guide future advances in agent memory for practical applications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.