MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Abstract
Memory has become a core component of modern LLM-based agents, enabling them to evolve from single-turn assistants into long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning. This risk becomes particularly important when historical memories are outdated, context-dependent, or inconsistent with stronger evidence available in the current task. Despite this emerging risk, existing memory benchmarks primarily evaluate whether memories are correctly stored, retrieved, or updated, while overlooking how retrieved memories influence downstream reasoning and decision-making. To bridge this gap, we propose MemSyco-Bench, a comprehensive benchmark for evaluating memory-induced sycophancy in agent systems. Rather than treating successful retrieval as sufficient, MemSyco-Bench measures when memory should influence a decision and how valid memory should be selected and used. Specifically, it covers five tasks that assess whether agents can reject memory as factual evidence, respect its applicable scope, resolve conflicts between memory and objective evidence, track memory updates, and use valid memory for personalization. In this way, MemSyco-Bench shifts memory evaluation from retrieval success toward the reliability of post-retrieval memory use. All related resources are collected for the community at https://anonymous.4open.science/r/MemSyco-Bench-7FCF.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.