acceptodds
Under review as a conference paper at ICLR 2027

Evolve to Learn: Gradient-Aligned Context Evolution for Model Self-Improvement

Abstract

Self-evolving language models can continuously acquire useful experience through test-time interaction, reflection, and prompt optimization, yet turning such experience into model itself to ensure self improvement remains challenging. On-policy self-distillation (OPSD) provides a natural mechanism for internalizing test-time information into model parameters, yet it largely assumes that the generated experience is worth learning. We show that this assumption does not always hold: experiments that improve test-time performance may provide limited or even harmful training signals. To better characterize which experience should be internalized, we systematically analyze diverse mutated experiments across multiple tasks and models, and find that gradient-based alignment is a more reliable indicator of downstream distillation gains than immediate test-time improvement. Based on this finding, we propose E2L, a self-evolution framework that iteratively generates, evaluates, and distills experiments according to their estimated training utility. E2L closes the loop between experience generation and parameter learning, enabling agents to autonomously produce experience suitable for persistent self-improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.