EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
Abstract
Large language model (LLM) agents perform strongly on many benchmarks, yet most evaluations assume static environments, while real-world agents must continually adapt to changing environments. To address this gap, we introduce EvoArena, a benchmark that models progressive environment evolution across terminal, software, and social domains, and EvoMem, a patch-based memory paradigm that records memory evolution as structured update histories. Current agents achieve only 39.6% average accuracy on EvoArena, while EvoMem improves EvoArena by 1.5% and standard benchmarks GAIA and LoCoMo by 6.5% and 3.3%, respectively. EvoMem further improves chain-level accuracy by 4% where success requires sustained reliability across evolving states, with analyses showing better preservation and use of evolving evidence. These results highlight the importance of explicitly modeling environmental evolution in both agent evaluation and memory. We will opensource the code and data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.