acceptodds
Under review as a conference paper at ICLR 2027

Choose Your Moment: Selective Evolution of LLM Agents via Latent Reasoning Dynamics

Abstract

Long-horizon LLM agents can continually improve by consolidating accumulated experience into their model parameters, yet indiscriminate updating is neither necessary nor always beneficial. Frequent updates incur substantial training costs and may overfit transient experience or interfere with previously acquired capabilities, whereas delayed updates can force agents to repeatedly reconstruct reasoning patterns that could have been internalized. This raises a key question for self-evolving agents: when is model evolution actually worthwhile? In this paper, we study this problem through the lens of latent reasoning dynamics and introduce Latch, a selective evolution framework that encodes latent reasoning trajectories and extracts deep-to-shallow reasoning corrections. We further characterize whether these corrections are locally realizable by a shared parameter change and remain consistent across past queries, providing evidence of whether the accumulated reasoning patterns can be jointly internalized. These signals are then used to predict the finite-horizon advantage of updating immediately over waiting, supervised by paired update–wait rollouts that measure downstream task performance and capability retention. For evolution steps predicted to be worthwhile, we perform on-policy self-distillation, using a frozen deeper reasoning path to supervise the current shallow model on student-generated prefixes while using experience replay to mitigate degradation of prior capabilities. Experiments across long-horizon benchmarks demonstrate consistent improvements in cumulative task utility and capability retention with reduced update overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.