acceptodds
Under review as a conference paper at ICLR 2027

SkillHelix: Coevolving Agent Skills via Decision and Reasoning Synergy

Abstract

Agent skills are natural-language documents that guide a model through a class of tasks. The model can improve these skills by examining its failures and rewriting the instructions. Such loops assign credit to entire rewrites. They reveal whether a rewrite helped, but not which edits contributed. Harmful edits are therefore accepted alongside helpful ones, and the optimizer skill guiding the rewrites cannot learn which kinds of edits work. Assigning credit to individual edits requires counterfactual evaluations, with up to runs for edits. Yet accuracy on the small batch used in each round rarely distinguishes one edit combination from another. We obtain this credit by pairing the black-box reasoning model with a small open-weight decision model. The reasoning model's outcomes ground the signal in its own behavior, identifying the questions an edit should fix and those it must preserve. The decision model provides a dense signal through forward passes that measure how each edit combination changes the likelihood of the gold answers on these questions. SkillHelix uses this signal in two ways. It selects edits for the task skill and refines the wording of their most influential span, retaining the change only if validation accuracy improves. It also assigns a Shapley value to every proposed edit and uses these credits to revise the optimizer skill, allowing the two skills to coevolve. With GPT-5.5 as the reasoning model, SkillHelix outperforms all baselines on six benchmarks, improving accuracy by an average of 3.8 points over the strongest baseline and by 12.9 points on LiveMath. Ablations show that edit-level decisions and coevolution both contribute to these gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.