acceptodds
Under review as a conference paper at ICLR 2027

SKILL-AS-CONTROLLER: A NEW PARADIGM FOR SKILL UTILIZATION IN LONG-HORIZON AGENTS

Abstract

Skills are crucial for LLM agents performing long-horizon tasks. However, when used only as prompt context, their effect on action selection remains implicit and suboptimal. Our analysis reveals a skill utilization gap: agents often violate skill rules even when relevant skills are provided, while the same model can often detect these violations afterward. We therefore propose **Skill-as-Controller**, a new paradigm that uses skills to explicitly control the action distribution. At each action step, Skill-as-Controller evaluates candidate actions based on the current state and skill rules, producing skill–action consistency scores. These scores are then combined with the base action distribution through a KL-regularized policy update, yielding a closed-form skill-controlled policy without model parameter updates. Experiments on three long-horizon benchmarks with diverse LLM backbones show consistent improvements in task success and interaction efficiency. On ALFWorld, Skill-as-Controller enables Qwen-3.5-4B to achieve a 92.5% success rate, exceeding the baseline by 58.9 points, while reducing the average number of interaction steps from 24.93 to 12.40. Further analysis shows that Skill-as-Controller reduces common failures, such as state loops and prerequisite violations, while the consistency scores make the effects of skills easier to inspect. These results establish Skill-as-Controller as an effective, efficient, and interpretable paradigm for skill utilization in long-horizon agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.