Unifying Text Memory and Code Memory for Self-Evolving Agents
Abstract
Self-evolving agents improve over time by distilling experience from past executions to facilitate future tasks. Existing systems represent such experience either as natural-language text injected into the agent context or as code that can be invoked via tools. However, the trade-offs between the two representations are poorly understood, and the choice between them is made a priori at the system level rather than based on the characteristics of each individual experience. We present the first controlled study that compares text memory and code memory under an identical task stream. The results show that the two memory forms are complementary in terms of construction cost, execution efficiency, and transfer robustness, such that neither representation alone is sufficient. Guided by these findings, we propose Metis, a self-evolving agent system that unifies text memory and code memory in a hierarchical manner. In particular, Metis organizes textual experience into execution plans, environment facts, and common pitfalls, and selectively crystallizes recurring plans into callable code tools. This design combines the good generality of text memory with the high execution efficiency of code memory and ensures that tool-generation cost is justified by repeated reuse. We evaluate Metis on three challenging benchmarks for interactive agents. Compared with a no-memory agent, SOTA baselines improve average accuracy by 9.5%–15.3%, while Metis improves accuracy by 26.9% and matches the most efficient baseline in execution cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.