Self-Hinting Language Models Enhance Reinforcement Learning
Abstract
Group Relative Policy Optimization (GRPO) often stalls on hard prompts under sparse terminal rewards: identical rewards within a rollout group collapse relative advantages and eliminate updates. We propose self-hint aligned GRPO with privileged supervision (SAGE), an on-policy method that uses training-only hints compressed from reference solutions to reshape the rollout distribution. For each prompt , the model samples a compact hint and generates a solution conditioned on , while retaining the same verifier reward . At each training step, the scheduler starts without a hint and increases hint strength only if a rollout group has no correct response. This makes informative groups more likely under finite sampling and can restore learning signals on hard prompts. At test time, we deploy the no-hint policy with . By refreshing self-hints as the learner changes, SAGE provides a better-calibrated curriculum than fixed hints from an initial policy or an external model. Across six benchmarks and three backbones, SAGE consistently outperforms GRPO, gaining up to 4.0 percentage points in average accuracy when trained on hard prompts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.