acceptodds
Under review as a conference paper at ICLR 2027

RecurLens: Tracing the Evidence for Recursive Self-Improvement

Abstract

When an agent improves after revising itself, the final task score does not reveal which part of the revision worked. The new method may never reach a successor, and a stronger descendant may still be worse at producing its next update. We introduce , an executable audit for fixed-weight agents that records proposed and installed policies, follows their use across three update opportunities, and tests inherited methods in a common recipient. The audit exposes separations that an endpoint score hides. Sol's Recursive descendant scores 5.38 points above its Frozen counterpart but 17.09 below its own initial state, with no method change. A supplied method reaches 136 successors in a later handoff condition, while failed qualification checks limit the utility inference. In four independent packages, its full downstream contrast falls from 77.72 to 21.98 points when the task policy written by the first update is removed. Thus task gains, actual inheritance, and later improvement utility require separate evidence; the recorded runs establish some links while leaving others unresolved. We will release the benchmark implementation and reproducibility materials after review.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.