acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Self-Evolving Agent Memory: Content, Writer, and Source

Abstract

Large language model agents can improve through interaction experience by storing and reusing external textual memory. However, there is little understanding of how such memory enables self-evolution and which aspects of its design actually drive downstream performance. Existing methods often simultaneously change what is stored, which model writes it, which experiences are used, and how the memory is maintained, making their contributions difficult to isolate. In this paper, we decompose the memory pipeline into controlled design dimensions and systematically study four common representations—raw trajectories, reflections, rules, and procedural skills—under a shared agent framework across six interactive benchmarks. We find that there is no universally best memory representation. Instead, performance depends on whether the abstraction of the stored experience matches the reusable structure of the environment. We further find that diversity does not consistently outperform repeated experience, whereas success-derived memory consistently outperforms failure-derived memory, with gaps of up to 38 percentage points. Together, our results shift the design of self-evolving agent memory from accumulating more experience and stacking additional mechanisms toward selecting reliable experience and matching its representation to the structure of the task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.