Environment-Staleness-Aware Replay for Language Agents
Abstract
Reusing past interactions can make reinforcement learning for large language model agents more efficient, but changes in tool reliability can render those interactions misleading even when the policy remains unchanged. We introduce Environment-Staleness-Aware Replay (ESR) to retain useful historical experience while accounting for changes in the environments that produced it. ESR executes the same actions from restored states in both environment versions and uses a regularized least-squares fit to estimate how the probabilities of tool outcomes have changed. It combines these estimates with policy corrections to reweight complete interaction sequences, excludes replay without sufficient execution evidence, and counts probing and restoration toward the interaction budget. On AppWorld, ESR improves mean post-change task-completion area under the learning curve by 2.44–4.37 points on a 0–100 scale over training with fresh interactions alone across three models and two test splits, with all methods sharing the same total interaction budget. These findings identify environment staleness as a distinct source of replay error and support evidence-based experience reuse in restorable environments with changing tool reliability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.