acceptodds
Under review as a conference paper at ICLR 2027

TEMPA: Temporally Extended Memory for Policy Adaptation

Abstract

Large language model (LLM) agents must learn from online experience to improve their behavior, even though their parameters remain frozen after deployment. Reusing complete episode trajectories as prompt context can incur additional token costs, while estimating individual action values from experience is challenging under sparse and delayed rewards. These challenges motivate policy adaptation through planning over action segments at an intermediate temporal granularity. We introduce TEMPA (Temporally Extended Memory for Policy Adaptation), a gradient-free framework that adapts an agent's policy over actor actions and executable memory segments. It partitions interaction at reward events and stores the resulting reward-to-reward segments as executable options. These segments are selected using empirical advantages. A router uses evidence from exploratory trials to choose between executing the selected segment and invoking the LLM actor. TEMPA reuses multi-step behaviors without separate credit estimates for constituent actions and reduces repeated actor inference during memory execution.We formulate memory exploitation as finite-horizon semi-Markov planning and establish a conditional sublinear bound on exploration loss relative to the incumbent policy. Experiments on Jericho and ScienceWorld show that TEMPA achieves higher average scores and stronger late-episode performance. Its total inference token consumption is roughly 56% lower than the mean across the seven baselines. Ablations support the contributions of reward-based segmentation, continuation values, adaptive routing, and persistent execution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.