acceptodds
Under review as a conference paper at ICLR 2027

Self-Generated Supervision for Self-Evolving Agents

Abstract

Long-horizon interactive agents often rely on task rewards, expert trajectories, or stronger teacher models to improve their performance. However, many complex, open-ended tasks lack verifiable outcomes and reliable external supervision, making it difficult for agents to improve through these approaches. Yet agents can actively interact with their environments, and the feedback from these interactions provides a natural source of learning signals. Building on this observation, we propose a simple framework for online agent self-improvement through on-policy self-distillation. At selected student-visited states, the current student explores alternative actions under a larger interaction budget, while an EMA teacher uses the resulting environment feedback to make a better-informed decision that is distilled into the zero-lookahead student. Our method generates supervision at individual decision steps, without relying on task rewards, explicit value estimates, or the outcomes of complete trajectories. As the student policy is updated, it generates new higher-budget teacher supervision, allowing both the student and teacher to improve over time. Experiments on ALFWorld show that, without using task rewards or expert supervision during training, the student's task success rate increases from 28.4% to 90.3%, while that of the higher-budget teacher increases from 76.9% to 96.3%. The trained student also surpasses the initial higher-budget teacher, demonstrating that agents can achieve sustained self-improvement through their own interactions with the environment. Our code and models will be publicly released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.