Skill Withdrawal in Self-Evolving Language Agents: A Causal Study of Non-Parametric Deskilling and Structural Dependency
Abstract
Large language model agents can acquire reusable behaviors without parameter updates by accumulating executable skills and persistent external state. Existing evaluations primarily measure performance while these resources remain available, so the effect of prolonged skill use on subsequent autonomous problem solving after skill withdrawal remains unclear. We define non-parametric deskilling as a counterfactual withdrawal gap: under an identical skill-free interface, a skill-exposed agent performs worse than a matched agent that has never received high-level skills, although both use the same frozen model, initial capability, task distribution, atomic tools, and evaluation budgets. We introduce a randomized longitudinal protocol with never-skilled, yoked-experience, information-matched, compute-matched, and state-reset controls to separate persistent behavioral changes from the immediate loss of external procedures, information, or computation. Versioned interventions on experience memory, retrieval statistics, planning templates, and controller state identify whether accumulated non-parametric state mediates withdrawal deficits. Exposure-dose randomization and post-withdrawal tracking characterize dependence, recovery sample complexity, and the area under the recovery-deficit curve. We further propose Interleaved Skill-Free Rehearsal (ISR), which hides the complete skill library during randomly selected exposure episodes and preserves trajectories of end-to-end task decomposition, atomic-tool use, execution, and verification within the same persistent state. We evaluate the framework across reproducible interactive environments and frozen model families using task success, constraint satisfaction, failure modes, state mediation, exposure dependence, and skill-free recovery. The framework supports controlled analysis and mitigation of persistent reliance on external skills while keeping model parameters fixed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.