DeltaSkill: Residual Skill Compilation and Evidence-Guided Control for Long-Horizon Terminal Agents
Abstract
Skills have enabled a new class of language-model agents that draw on modular procedures and execution strategies to solve complex tasks. However, poorly adapted skill use can prolong task execution and waste tokens. In this work, we study adaptive skill use in language-model agents, combining capability-aware residual skill compilation with evidence-guided execution control. We call this framework DeltaSkill. We evaluate DeltaSkill against seven baselines on three long-horizon benchmark tasks, comparing execution time, task scores, and token consumption. Our findings reveal that lower execution costs can accompany improved or near-best solution quality. These benefits include improved reward on Generals, near-best quality on Spot with 57.2% fewer execution tokens and 84.7% less recorded wall time than Structured Local, and complete artifact delivery on DuckDB, where the seven baseline executions remain incomplete. To our knowledge, this is the first study to couple receiver-conditioned residual skill compilation with the retention of validated task artifacts and evidence-guided continuation for long-horizon terminal agents. Finally, we outline directions for extending DeltaSkill to continual capability profiling and cross-task skill adaptation, toward more efficient and generalizable agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.