CAVRA: Causal Adaptive Verified Remasking for Self-Evolving Agents
Abstract
Large language model agents are increasingly deployed as long-horizon controllers, yet they are almost always served with frozen weights, so any lasting improvement must live in the external textual policy—skills, plans, and memory—rather than in the parameters. Turning a failed trajectory into a policy that transfers remains obstructed by two protocol gaps: a teacher that locates the defect either consults privileged information that vanishes at deployment or else rewrites the entire episode, and a skill written from the remasked instance itself contaminates later tasks. We therefore propose CAVRA, a remask–recover framework in which the same frozen model is both the student that acts and the teacher that localizes failure. Causal remask attribution spends a bounded mask budget on the high-credit control span, reading the world's own stall signature rather than a privileged judge. Local recover then inpaints only those spans and scores each fill with a progress-aligned probe confined to the error region, so the official first-pass verdict is never overwritten. Utility-latched precipitation crystallizes a skill from remask kinds and failure reasons—or from already-corrected local fixes—after instance facts have been stripped, and family-conditioned routing injects only the matching skeleton. Extensive experiments, including a matched-backbone comparison against ten baselines, module ablations, and adaptation-cost measurements, show that CAVRA improves interactive control substantially, leading the strongest baseline by 40.6% relative success on ALFWorld.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.