ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning
Abstract
Reinforcement learning with verifiable rewards has become the dominant paradigm for mathematical reasoning in language models, yet most methods treat each problem as a fresh episode and retain discovered strategies only implicitly through policy updates. Abstracting reusable techniques and recognizing when to apply them, the central mechanism of mathematical expertise, therefore stays outside training. Unlike the executable procedures stored in skill libraries for tool-use and embodied agents, a mathematical skill is a reasoning pattern embedded in a solution trace, whose applicability must be inferred from the problem and whose worth is defined only against the same policy solving it unaided. Abstracting and selecting such skills are therefore reasoning tasks that existing designs leave outside policy optimization. We introduce ARISE (Agent Reasoning with Intrinsic Skill Evolution), a hierarchical reinforcement learning framework in which one shared policy serves as the high-level option policy that selects from a persistent skill library and as the low-level intra-option policy that generates solutions conditioned on the selected skill. The policy scores library entries through its own conditional log-likelihood and concludes new skills from ground-truth traces through an inference step. With one option per query, the option-critic gradient reduces to two terms under one group-relative advantage, and generation feedback reaches selection through the tangent kernel between skill documents and rollouts, so a single policy gradient optimizes both levels. The three-level trajectory reward, which separates skill-aided from unaided correct solutions, enters the advantage, while a counterfactual skill quality level governs library maintenance, so reasoning ability and library quality co-evolve. Across two instruction-tuned base models and seven competition and Olympiad-level benchmarks, ARISE outperforms GRPO-family algorithms and memory- and skill-augmented baselines, with the largest gains on out-of-distribution tasks, and the gains persist on a third backbone at 8B scale and in code generation. The implementation is available at [ARISE Codebase](https://anonymous.4open.science/r/ARISE_ICLR-69D2).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.