acceptodds
Under review as a conference paper at ICLR 2027

Skill2Task: Learning from Skills through Graph-Guided Task Synthesis for Agentic RL

Abstract

Large language model (LLM) agents can reuse past experience during reinforcement learning (RL) through natural-language skills distilled from their trajectories on training tasks. However, we uncover a bottleneck in SkillRL: generalization can stall in later training stages. Skills distilled from individual task trajectories may also contain overly broad, conflicting, or incorrect recommendations. These challenges motivate broadening the agent’s training experience to support further policy improvement and provide opportunities to test potentially unreliable recommendations through execution feedback. We therefore introduce a framework that uses skills to specify and construct training tasks with varied goals and conditions. To this end, we organize skills into a graph that captures their recommendations and applicability conditions, together with relations that describe how skills can be combined and when their recommendations conflict. The graph guides task synthesis in existing environments, with verifiable success criteria and checks for task relevance and solvability. Starting from a SkillRL checkpoint, we incorporate the resulting tasks into continued RL, using environment feedback to train the agent to combine skills and decide when to follow, adapt, or reject their recommendations. On ALFWorld, Qwen3-4B-Instruct trained with tasks synthesized by our framework achieves higher average success rates than SkillRL on both seen and unseen tasks, with gains of 8.1 percentage points on unseen Pick&Place tasks and 11.1 points on unseen Examine tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.