acceptodds
Under review as a conference paper at ICLR 2027

SkillBackprop: Two-Stage Textual Backpropagation for Skill Optimization

Abstract

Optimizing skills for frozen language models requires more than identifying failures. It also requires assigning task loss appropriately and determining how the skill should be revised. Unlike differentiable model parameters, skill instructions do not admit analytic derivatives that can directly guide these decisions. To provide a backward learning pathway for skills, we introduce SkillBackprop, a two-stage textual backpropagation framework for structured skills. SkillBackprop separates backward learning into failure localization and loss-reduction estimation. The first stage aggregates execution evidence to trace failed outcomes to the responsible parts of the skill. The second measures how candidate revisions affect validation loss and uses the resulting loss-reduction signals to guide skill updates. Together, the two stages provide a backward pathway from task-level loss to targeted modification. We evaluate SkillBackprop on SpreadsheetBench, OfficeQA, SearchQA, and DeepPlanning-Shopping, where it achieves scores of , , , and , respectively. Across all benchmark and model settings, SkillBackprop matches or outperforms the strongest baseline, with gains of up to percentage points. Ablations further show complementary benefits from failure localization and loss-reduction estimation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.