acceptodds
Under review as a conference paper at ICLR 2027

Learning Without Remembering: Agent Self-Improvement through Experience Compilation

Abstract

Synthetic interaction data provide a scalable resource for training tool-using agents, yet successful demonstrations may not meet the learning needs of the current student. We investigate how feedback accumulated during generation and learning can be translated into requirements for constructing future training data. We introduce Learning Without Remembering (LWR), an agent learning framework that compiles stage-level experience into synthesis policies. LWR jointly analyzes generation successes and failures alongside changes in student performance and behavior before and after training, distilling synthesis guidelines with explicit applicability conditions that can be reused across examples. These guidelines are translated into generation instructions, workflow allocation, and interaction plans. With the generator and controller weights held fixed, the updated policy guides the generation of complete interaction trajectories, which undergo environment verification before being used for continued supervised fine-tuning of the student. Subsequent generation and learning feedback supports further policy revision. Past experience is thus consolidated into generation requirements and incorporated into student parameter learning through new training data, without retrieving past cross-task trajectories at test time. Experiments on and AppWorld with two student model sizes demonstrate improved task completion. Compared with fixed synthesis, LWR improves macro-average success rates by 5.00 and 6.67 percentage points for the 8B and 30B students, respectively, and yields additional gains on most AppWorld metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.