acceptodds
Under review as a conference paper at ICLR 2027

Skill in the Loop: Explicit Skill Declaration and Attribution for Agent Training

Abstract

Agents increasingly rely on reusable skills distilled from historical experience to solve long-horizon decision-making tasks. However, existing methods primarily provide skills as contextual information, leaving the agent's skill selection implicit and offering limited mechanisms to assess and improve how skills are used. We propose Skill-Conditioned Policy Optimization (SCPO), a closed-loop framework for explicit skill use and continual skill evolution. Before each action, the agent explicitly declares the skill guiding its decision. A retrospective mechanism then evaluates the effectiveness of skill usage over multi-step interactions and provides fine-grained feedback, which jointly optimizes the policy via GRPO and updates the skill bank based on observed utility. Experiments on ALFWorld and WebShop show that SCPO achieves state-of-the-art performance and consistently improves both skill utilization and skill-bank quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.