acceptodds
Under review as a conference paper at ICLR 2027

Talentrl: Strategy-Guided Reinforcement Learning for Hierarchical Skill Routing

Abstract

Natural-language skills let large language model (LLM) agents reuse procedures across tasks, yet most systems retrieve individual Skills directly from a flat SkillBank without explicitly modeling the high-level strategy that governs their selection and use. We introduce TalentRL, a strategy-guided reinforcement learning framework for hierarchical routing over procedural memory. Coarse-grained Talents summarize recurring skill requirements to guide routing, while fine-grained Skill cards provide actionable procedures. For each task, the policy selects a Talent to guide execution and softly prioritize its associated Skills during similarity-based retrieval. However, because task outcomes depend jointly on Talent selection, retrieved Skill context, and downstream execution, terminal rewards cannot directly credit the routing decision. TalentRL therefore trains Talent routing by comparing rewards across matched intervention rollouts, while optimizing downstream execution with group-relative advantages. Across seven open-domain and multi-hop question-answering benchmarks, Talent-aware supervised fine-tuning improves average success rate by 7.53% for Qwen2.5-7B and 6.71% for Qwen3-8B over their zero-shot baselines. Reinforcement learning further improves the Talent-aware SFT checkpoint, enabling TalentRL to achieve performance competitive with state-of-the-art SkillRL across the seven benchmarks. These gains extend to the agentic benchmarks ALFWorld and WebShop, where TalentRL consistently improves over SkillRL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.