acceptodds
Under review as a conference paper at ICLR 2027

Policy-Relative Frontier Learning for Self-Evolving Language Models

Abstract

As language agents are deployed across increasingly diverse tasks, self-evolution offers a practical way to adapt them through their own experience under limited budgets and strict data privacy requirements. Yet many real-world domains provide only outcome feedback, leaving the policy's current capability bottlenecks unclear. Each round of self-evolution therefore raises a central question: what should the policy learn next, and how far should supervision extend? We propose **Policy-Relative Frontier Learning**, a framework that uses progress differences among the policy's own rollouts to determine the boundaries and granularity of supervision. Using the shortest successful rollout as a reference, we identify the furthest shared progress of each failed rollout as its relative frontier. Each failed rollout learns from its own frontier to the next frontier in the group, or to task completion if no further frontier exists. By focusing on local progress, this design recognizes that even a rollout that ultimately fails can contain useful segments that help another attempt cross its relative frontier. We therefore also select segments from failed rollouts, expanding supervision beyond successful trajectories. We fine-tune the policy on the selected segments and reconstruct learning targets from fresh experience after each update. supervision from failed rollouts. Experiments on ALFWorld and WebShop show that our method outperforms baselines with fewer rounds of evolution. Ablations reveal that the effectiveness of fixed supervision granularity varies across tasks, highlighting the value of policy-relative learning intervals derived from the policy’s own progress. Code is available at https://anonymous.4open.science/r/relative-frontier-evolve-657F/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.