Learning from Hints, Forgetting Less: Reward-Directed Transfer for Continual Post-Training
Abstract
Training-time hints can help models solve difficult tasks, creating successful experience from which to learn. In continual post-training, the challenge is to turn these assisted successes into independent capabilities while retaining earlier ones. However, solely imitating hint-assisted success can pull the model toward particular answers or hint-dependent predictions, introducing behavioral changes that may interfere with previously acquired capabilities. We introduce Hint-Gap Weighted Transfer (HiGHT), which combines online reinforcement learning (RL) with selective transfer from hinted rollouts. Its central signal is the *hint gap*: the signed change in a sampled token's probability induced by the hint at the same response prefix. Positive trajectory advantages select reward-supported experience, while the hint gap modulates the direction and strength of each token's transfer contribution. Raw-prompt reinforcement learning continues to optimize task reward without hints. Across five real-world vision-language domains, HiGHT improves mean final accuracy from 49.94% to 55.12% and backward transfer from to percentage points over hinted-policy distillation combined with online RL (Distill + RL). Controlled local updates show better early-stage retention at matched update norms. These results support learning from hint-assisted experience with less forgetting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.