ROSE: Rollout-Free Skill Evolution via Static Task–Skill Analysis
Abstract
Skill evolution offers a practical way to improve agent performance without updating model weights. Existing methods often rely on repeated task rollouts and evaluator feedback to guide skill revision, making skill construction costly and time-consuming. Yet procedural deficiencies, such as formatting errors and omitted steps, can often be identified from task requirements before execution. Motivated by this observation, we introduce ROSE, a rollout-free framework for efficient agent skill evolution inspired by static program analysis. ROSE organizes an initial skill, training-task specifications, and visible artifact metadata into a shared code-style intermediate representation. It then audits this representation for unmet requirements, contract mismatches, and predictable execution risks, and uses the findings to guide localized skill revisions while preserving useful procedures. Construction requires only three LLM calls, without task rollouts or evaluator feedback. The resulting skill is frozen and reused across held-out tasks without test-time adaptation. Across SpreadsheetBench, LiveMathematicianBench, DocVQA, and OfficeQA with four agent models, ROSE achieves competitive performance, obtaining a macro-average score of 58.5%, compared with 56.8% for SkillOpt. On matched SpreadsheetBench construction instances, ROSE uses 9,183 tokens and completes construction in 32.7 seconds, compared with 3.86 million tokens and approximately 100 minutes for SkillOpt. These results demonstrate that static task evidence can support competitive held-out performance with substantially reduced skill-construction cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.