acceptodds
Under review as a conference paper at ICLR 2027

Self-Hinting Language Models Enhance Reinforcement Learning

Abstract

Group Relative Policy Optimization (GRPO) often stalls on hard prompts under sparse terminal rewards: identical rewards within a rollout group collapse relative advantages and eliminate updates. We propose self-hint aligned GRPO with privileged supervision (SAGE), an on-policy method that uses training-only hints compressed from reference solutions to reshape the rollout distribution. For each prompt , the model samples a compact hint and generates a solution conditioned on , while retaining the same verifier reward . At each training step, the scheduler starts without a hint and increases hint strength only if a rollout group has no correct response. This makes informative groups more likely under finite sampling and can restore learning signals on hard prompts. At test time, we deploy the no-hint policy with . By refreshing self-hints as the learner changes, SAGE provides a better-calibrated curriculum than fixed hints from an initial policy or an external model. Across six benchmarks and three backbones, SAGE consistently outperforms GRPO, gaining up to 4.0 percentage points in average accuracy when trained on hard prompts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.