acceptodds
Under review as a conference paper at ICLR 2027

Causal Identification of Endogenous Feedback Amplification in Self-Evolving Agent Skills

Abstract

Large language model agents convert execution trajectories into persistent skills, creating an endogenous loop in which the current skill shapes the evidence that drives its subsequent evolution. Following a controlled skill shift, a single evolving run cannot distinguish among static retention, repeated evidence exposure, stochastic drift, and amplification through the skill-to-trajectory-to-skill pathway. We introduce Paired Counterfactual Skill Dynamics (PCSD), a causal framework that identifies this endogenous feedback effect. PCSD synchronizes two branches: a closed-loop branch, in which each updated skill generates evidence for its successor, and an open-loop shadow branch, in which the initially shifted skill is held fixed and always generates evolution evidence, while updated shadow skills are evaluated but cannot influence subsequent trajectories. The paired difference isolates the effect mediated by skill-induced changes in trajectory distributions while holding the initial shift, task stream, update budget, and evolver calls fixed. Information-matched replay distinguishes feedback from equivalent evidence exposure, while a frozen-skill branch estimates generation drift and measurement noise. PCSD also decomposes local loop gain into behavioral and evolution responses. We vary displacement, update intensity, history window, deployment policy, and clean-data anchoring to characterize stability boundaries, hysteresis, recovery, and path dependence. Deterministic probes, paired controls, and explicit rejection criteria enable evaluation across isolated task domains and structurally distinct evolution regimes while separating endogenous amplification from ordinary retention and finite-horizon variation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.