MemEnv: Evaluating Multi-dimensional Long-Term Memory via Graph-Grounded Dynamic Interaction
Abstract
Long-term memory is essential for LLMs to function as coherent and personalized assistants. Yet, current evaluations often rely on static histories that fail to capture real-world conversational dynamics or only assess narrow memory traits. We introduce MemEnv, a dynamic benchmark designed to assess multiple dimensions of long-term memory across sessions. In MemEnv, each memory dimension is grounded in a scalable, predefined graph. A user simulator leverages this structure to orchestrate conversations, seamlessly interleaving predefined utterances with open-ended dialogue before querying the model. To better mimic the complicated nature of real-world interactions, we further present MemEnv+, which integrates multiple memory dimensions into individual conversations. Extensive experiments reveal that current LLMs and memory agents still lack well-rounded memory capabilities. While memory mechanisms boost overall performance, they introduce clear trade-offs across different memory dimensions. MemEnv+ exacerbates these challenges, especially when models are required to synthesize dispersed evidence. Furthermore, in-conversation backbone-model switching reveals a gap between on-policy and off-policy settings, with off-policy achieving larger gains on information-focused tasks regardless of the preceding model's strength. These findings highlight the need for versatile memory systems that can balance different dimensions and perform reliably in dynamic environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.