Decision-Centric Reinforcement Learning over Promising Tokens for Large Language Model Reasoning
Abstract
Reinforcement learning (RL) for large language models (LLMs) is commonly formulated as token-level policy optimization over the full vocabulary. This fixed action space is tens of thousands of tokens wide, although only a small state-dependent subset typically represents plausible next reasoning moves. Existing decoding strategies such as Top-k or nucleus sampling exploit this structure during rollout, but standard RL objectives still compute policy ratios and gradients with respect to the full-vocabulary policy, creating a mismatch between the space where trajectories are collected and the space where the policy is optimized. In this work, we study LLM generation as decision-making over a fixed global action space with a dynamic local decision set. We introduce Reinforcement Learning with Promising Tokens (RLPT), which uses the behavior policy’s own distribution to construct a promising token set at each state and defines a masked policy consistently for both rollout and optimization. This simple policy reparameterization turns token-level RL from full-vocabulary fitting into action selection among plausible reasoning candidates. Across mathematical reasoning, code, and domain QA tasks, RLPT improves GRPO and DAPO, with multi-seed results on Math-17k, MATH500, AIME24, and AIME25 showing improved accuracy and reduced training variability. Controlled ablations further show that the gains come from aligning rollout and optimization supports, rather than from decoding-time truncation alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.