acceptodds
Under review as a conference paper at ICLR 2027

Memory-Driven Self-Improving LLM Agents for Sequential Decision-Making

Abstract

Large language models (LLMs) have emerged as effective action policies for sequential decision-making (SDM) tasks due to their extensive prior knowledge. However, this broad yet general knowledge is often insufficient for long-term decision-making tasks with limited task-related experience, making it challenging to efficiently adapt LLMs to specific SDM tasks, especially for generalizing to unseen tasks. To address this challenge, we propose a memory-driven self-improvement framework for LLM agents that combines the general prior knowledge of LLMs with a compact memory of domain-specific experiences. The framework consists of two mutually reinforcing components: memory-driven value estimation and LLM prior refinement. Specifically, the memory stores state-action experiences with Q-values grounded in environmental feedback, enabling retrieval-based, non-parametric value estimation while providing valuable experiences for refining the LLM prior. The refined LLM prior, in turn, generates higher-quality trajectories that further enrich the memory, forming a self-improvement loop. Experiments on multi-step agentic and mathematical reasoning tasks demonstrate that our approach improves sample efficiency and generalization to unseen tasks, substantially outperforming traditional RL and LLM-based baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.