Learning to Hint for Reinforcement Learning
Abstract
Group Relative Policy Optimization (GRPO) often suffers from advantage collapse in reinforcement learning with verifiable rewards: when all rollouts in a group receive the same reward, their relative advantages vanish. Hint-based methods address all-incorrect groups by adding guidance that induces mixed outcomes and recovers a learning signal. However, creating a learning signal under the hinted input does not necessarily improve the no-hint policy used at test time. To this end, we propose **Hi**nt **L**earning for Reinforcement **L**earning (HiLL), a framework that learns to generate hints by considering both signal creation and signal transfer. HiLL trains a hinter alongside the reasoner, generating hints conditioned on the question, an incorrect rollout from the current reasoner, and a reference solution available only during training. We introduce *hint reliance*, which measures how strongly correct hinted trajectories depend on the hint by comparing their likelihoods with and without it. We derive a theoretical transfer bound relating hint reliance to hinted and no-hint success probabilities. Motivated by this result, we construct a transfer-weighted hinter reward that favors hints that recover informative GRPO groups while penalizing high hint reliance. Experiments across eight benchmarks with two backbone LLMs show that HiLL outperforms GRPO and hint-based baselines. Ablations further show that transfer weighting reduces hint reliance and improves no-hint performance, demonstrating the value of transfer-aware hint learning for RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.