acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical-R1: Hierarchical Data Selection for RL Post-training of LLMs

Abstract

Reinforcement learning (RL) post-training has emerged as an effective paradigm for improving large language models (LLMs), particularly on reasoning and planning tasks. Recent approaches leverage token entropy to identify informative tokens for RL optimization. However, existing token-selection methods typically treat high-entropy tokens uniformly, overlooking their heterogeneous effects across reasoning trajectories and potentially limiting effective exploration. Motivated by a systematic analysis of high-entropy tokens, we introduce Hierarchical-R1, a hierarchical data selection framework that integrates sample-, trajectory-, and token-level signals for RL post-training. Specifically, trajectory-guided token selection differentiates high-entropy tokens from positive and negative trajectories to better preserve policy entropy and sustain exploration, while token-guided sample selection prioritizes samples likely to contain informative tokens, improving training efficiency. Experiments on models with different sizes across six mathematical reasoning benchmarks show that trajectory-guided token selection consistently improves accuracy over the evaluated baselines, while incorporating token-guided sample selection further reduces training cost with comparable average accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.