Self-Evolving Agents through Privileged Retrospective Distillation
Abstract
Self-evolving agents aim to improve from their own experience rather than remaining fixed once deployed. In practice, however, naturally collected agent trajectories often lack ground-truth solutions or verifier-derived rewards, making them difficult to learn from. We observe that these trajectories become substantially more informative in retrospect. When the agent takes an action, it knows only what has happened so far, but looking back after the trajectory ends, the same action can be judged with knowledge the agent lacked at that moment, including how the environment responded afterward and how related tasks in its accumulated experience succeeded or failed. Based on this observation, we introduce Retrospective Self-Distillation (RSD), a multi-round self-evolution framework in which the same model, given privileged retrospective context built from its own experience, acts as a teacher that re-scores its own actions, and the resulting distribution is distilled back into the policy through on-policy self-distillation. The privileged context combines cross-task precedents and lessons with turn-level reflection and future hindsight, allowing both successful and imperfect trajectories to provide dense supervision. RSD is also built to make this supervision improve alongside the agent. Every round, precedents and lessons are re-derived from the experience bank enlarged with the current policy's trajectories, and reflection and hindsight are computed fresh for each new trajectory, so the teacher always scores the policy against its most recent experience. We further introduce frontier-guided context-policy evolution, in which a stronger frontier model proposes edits to the context construction rules, for example how many precedents to retrieve or which components to include. Across ALFWorld, ScienceWorld, and SearchQA with Qwen3-8B and 14B, RSD improves performance by up to 59.7% relative over the base model, and frontier-guided context-policy evolution adds up to a further 44.7% relative improvement over a fixed context policy. These results show that agents can convert naturally accumulated experience into persistent model capability without ground-truth trajectories or verifier-derived rewards for the core policy update.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.