acceptodds
Under review as a conference paper at ICLR 2027

Environment-Staleness-Aware Replay for Language Agents

Abstract

Reusing past interactions can make reinforcement learning for large language model agents more efficient, but changes in tool reliability can render those interactions misleading even when the policy remains unchanged. We introduce Environment-Staleness-Aware Replay (ESR) to retain useful historical experience while accounting for changes in the environments that produced it. ESR executes the same actions from restored states in both environment versions and uses a regularized least-squares fit to estimate how the probabilities of tool outcomes have changed. It combines these estimates with policy corrections to reweight complete interaction sequences, excludes replay without sufficient execution evidence, and counts probing and restoration toward the interaction budget. On AppWorld, ESR improves mean post-change task-completion area under the learning curve by 2.44–4.37 points on a 0–100 scale over training with fresh interactions alone across three models and two test splits, with all methods sharing the same total interaction budget. These findings identify environment staleness as a distinct source of replay error and support evidence-based experience reuse in restorable environments with changing tool reliability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.