Beyond Memorization: Benchmarking Memory Utilization in Conversational LLM Agents
Abstract
Large Language Models (LLMs) agents increasingly rely on persistent cross-session memory to support long-horizon and personalized tasks. To understand how memory is actually needed in real-world use, we annotated WildChat and identified 21,382 memory-dependent queries, of which 83.7% do not ask about memory content directly but invoke it while solving downstream tasks. Existing memory benchmarks almost exclusively evaluate the former: whether memory can be accurately recalled, leaving the latter untested. Drawing on Bloom's taxonomy, we identify three higher-order memory abilities beyond basic recall: Comprehension of memory revealed in dialogue, Apply of memory to downstream tasks, and Evaluation of whether memory should be invoked or used to correct erroneous queries. We introduce MUSE-Bench, which operationalizes these abilities into six utilization scenarios under a need-first construction pipeline. Built upon 100 carefully constructed user profiles spanning 12 life domains, MUSE-Bench contains 3,840 queries across six evaluation tasks. Notably, the Evaluation level introduces two opposing failure modes: Interference Memory, where irrelevant or misleading memory must be suppressed; and Counterfactual Query, where memory must be invoked to correct factual errors in the query, targeting the deployment risks of memory overuse and sycophancy. We evaluate 14 frontier LLMs and 20 memory baselines on MUSE-Bench, and identify two key findings: i) Frontier LLMs perform reasonably at Comprehension but show substantial gaps at Apply and Evaluation level; ii) Existing memory techniques offer limited improvement in memory utilization over conventional RAG methods. These results suggest that the next frontier of agent memory is not memorizing more, but using more wisely.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.