Agent Skills Under Change: A Lifecycle Benchmark and Coordinate-Guided Updates
Abstract
As agents accumulate reusable Skills, their capabilities depend on artifacts that must evolve with changing facts, tools, and policies. Yet an update that fixes today's failure can silently break working behavior and propagate errors into future tasks. We introduce LifecycleSkillBench, a lifecycle evaluation protocol and benchmark that pairs reproducible environmental changes with affected and protected tasks and tracks recovery, retention, and error propagation across updates. We also propose CoordinateClosure, a coordinate-guided update framework spanning four dimensions: facts, interfaces, execution boundaries, and policies. It localizes edits, checks dependencies, and combines public repair feedback with execution replay to decide whether to adopt an update. In controlled development experiments with matched public inputs, model-call schedules, and output caps, coordinate-structured patches improve first-pass state-check acceptance from 68.8% to 98.4% for Qwen3.7-Plus and from 76.6% to 85.9% for GLM-5.2. In a separate study of four 16-step maintenance streams, the complete controller, including repair and checkpoint restoration, passes the evaluated state and safety checks in 64/64 states, compared with 56/64 for whole-state rewriting. Recorded input and output tokens fall from 958,833 to 42,225, a 95.6% reduction. These experiments show how lifecycle evaluation exposes error persistence across updates, while coordinate-structured patches improve first-pass acceptance and the full CoordinateClosure controller reduces recorded token use relative to whole-state rewriting in the tested maintenance settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.