Mem-Behave: Behavioral Stress Testing of Agent Memory Systems
Abstract
The growing adoption of persistent language-model agents makes reliable long-term memory increasingly important. Yet existing benchmarks typically evaluate memory at a single difficulty level, leaving unclear how performance changes as a capability is placed under increasing stress. We introduce Mem-Behave, a benchmark for stress-testing persistent agent memory across four capabilities: plasticity, interference robustness, contextual integrity, and forgetting. Each capability is paired with a stressor that can intensify over long-term use, namely reinforcement of outdated information, near-miss competing memories, ambiguity in disclosure context, and entanglement of information that should be forgotten. Built from long multi-session histories, Mem-Behave contains 767 scenarios each evaluated at three stress levels across four memory systems yielding 3,508 probes in total. Our results show that performance declines with increasing stress across all four capabilities and memory systems, but the magnitude and onset of degradation vary substantially across systems and capabilities. Further diagnostic analysis shows that many downstream failures are already apparent in the memories retrieved by the system. Together, these findings demonstrate that single-difficulty evaluation can obscure important differences in memory reliability, motivating stress-based evaluation of persistent memory systems. We will publicly release Mem-Behave upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.