acceptodds
Under review as a conference paper at ICLR 2027

Learner-Conditioned Experience Rewriting for Language Model Agents

Abstract

A successful agent trajectory demonstrates a solution, but does not identify which parts would improve the current learner. Failed trajectories can likewise contain useful actions. We introduce Learner-Conditioned Experience Rewriting (LER), a framework for converting both kinds of experience into supervision using three distinct signals: satisfaction of a checkable requirement, the learner's unaided reliability on that requirement, and the utility of a candidate update. A shared language model audits experience with training-only references and checkpoint-specific behavioral records, proposes local corrections, and learns from complete action turns admitted by explicit environment checks. Progress-conditioned replay allocates training within this verified pool. To train the rewriting policy, LER compares verified plans through equal-budget temporary updates from a common checkpoint and uses comparisons that pass relative-advantage and positive-gain filters as preference labels. The resulting objective targets measured short-horizon learning utility without differentiating through the updates. Explicit information boundaries and behavior-level measurements connect the construction, allocation, and evaluation of supervision. This formulation separates immediate correction quality from training utility, making subsequent unaided execution the objective of experience rewriting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.