Memento: Continual Learning of LLM Agents via Stateful Reflective Memory
Abstract
We present a unified method for continual learning of large language model (LLM) agents that fully eliminates the need to fine-tune the underlying model. Existing paradigms are either rigid (handcrafted reflection workflows) or computationally costly (gradient updates of LLM parameters), and neither scales to long-horizon tasks where environments drift faster than models can be retrained. We argue that the right primitive is reflection: revisiting past experience to adjust future action selection. We formalise this primitive as a Stateful Reflective Decision Process (SRDP), in which the agent alternately writes new experiences to an episodic memory and reads relevant cases to guide decisions. By augmenting the state space with the memory, we lift the SRDP to a Reflected Markov Decision Process amenable to classical control analysis. We then adapt KL-regularised policy iteration to this memory setting as Read–Write Reflective Learning, integrating Parzen-window retrieval with closed-form soft updates, and prove (i) convergence under bounded rewards and a fixed memory, (ii) two-time-scale convergence when memory evolves slowly, and (iii) asymptotic optimality as the memory densely covers the state space. We instantiate this method as a planner–executor Deep Research agent, Memento, with both non-parametric and parametric memory variants. Empirically, Memento statistically ties the strongest RL-trained baseline on the seven-dataset DeepResearcher suite without finetuning LLMs, and wins the average on five multimodal benchmarks. Furthermore, in end-to-end system-level comparison, Memento reaches Top-1 on the GAIA validation leaderboard, and surpasses prior agents on SimpleQA, HLE, and UIS-QA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.