Learner-Conditioned Experience Rewriting for Language Model Agents
Abstract
A successful agent trajectory demonstrates a solution, but does not identify which parts would improve the current learner. Failed trajectories can likewise contain useful actions. We introduce Learner-Conditioned Experience Rewriting (LER), a framework for converting both kinds of experience into supervision using three distinct signals: satisfaction of a checkable requirement, the learner's unaided reliability on that requirement, and the utility of a candidate update. A shared language model audits experience with training-only references and checkpoint-specific behavioral records, proposes local corrections, and learns from complete action turns admitted by explicit environment checks. Progress-conditioned replay allocates training within this verified pool. To train the rewriting policy, LER compares verified plans through equal-budget temporary updates from a common checkpoint and uses comparisons that pass relative-advantage and positive-gain filters as preference labels. The resulting objective targets measured short-horizon learning utility without differentiating through the updates. Explicit information boundaries and behavior-level measurements connect the construction, allocation, and evaluation of supervision. This formulation separates immediate correction quality from training utility, making subsequent unaided execution the objective of experience rewriting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.