Continual Policy Programming for LLM Agents: Learning Executable Policies from Online Experience
Abstract
LLM agents can solve increasingly complex long-horizon tasks, while continually improving from ongoing interaction remains an important challenge. We introduce Continual Policy Programming, a framework that enables a frozen LLM agent to learn from fine-grained online experience by updating a persistent executable policy. Policy updates take effect during the current task and persist across future tasks, allowing newly acquired knowledge to accumulate over time. Across ALFWorld, Webshop, MiniGrid, and Appworld, Continual Policy Programming enables agents to recover from failures online, reuse newly acquired behaviors across subsequent tasks, and accumulate capabilities over time. Our results show that persistent executable policies provide an effective way to turn streaming interaction experience into cumulative behavioral improvement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.