acceptodds
Under review as a conference paper at ICLR 2027

Continual Policy Programming for LLM Agents: Learning Executable Policies from Online Experience

Abstract

LLM agents can solve increasingly complex long-horizon tasks, while continually improving from ongoing interaction remains an important challenge. We introduce Continual Policy Programming, a framework that enables a frozen LLM agent to learn from fine-grained online experience by updating a persistent executable policy. Policy updates take effect during the current task and persist across future tasks, allowing newly acquired knowledge to accumulate over time. Across ALFWorld, Webshop, MiniGrid, and Appworld, Continual Policy Programming enables agents to recover from failures online, reuse newly acquired behaviors across subsequent tasks, and accumulate capabilities over time. Our results show that persistent executable policies provide an effective way to turn streaming interaction experience into cumulative behavioral improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.