From Experience to Policy: Structured Memory for Continual Learning in LLM Agents
Abstract
Large Language Model (LLM) agents increasingly rely on external memory for long-horizon reasoning. However, existing mechanisms primarily store passive interaction histories, which support retrieval but lack explicit reusable decision policies. Consequently, agents struggle to accumulate and refine experience across tasks. We propose Structured Policy Memory (SPM), a continual learning framework that transforms interaction trajectories into explicit policy knowledge organized as a structured graph. SPM establishes a closed-loop process where the agent retrieves relevant policies to guide decisions and validates newly extracted knowledge to update the memory after each interaction. Extensive experiments on LoCoMo, LongMemEval, StrategyQA, and ALFWorld demonstrate that SPM consistently outperforms representative prompting, retrieval, and memory-augmented baselines while maintaining computational efficiency. Specifically, SPM achieves 33.25% overall Semantic Similarity on LoCoMo (the best among all compared methods), exhibits the largest continual learning gain of +13.17pp across 1500 tasks, and attains 95.78% Action Match on ALFWorld. These results demonstrate effective experience reuse and sustained performance improvement, providing a practical foundation for continual learning in LLM agents. Our code is available at https://github.com/linda-2-hh/SPM.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.