Simple Self-distillation through Error Verbalization
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has become the de facto approach to enhancing model performance on tasks such as coding and mathematical reasoning. A common feature among popular RL algorithms, such as GRPO, is to assign every token of a failed rollout the same negative advantage, penalizing sound and unsound reasoning steps alike. *Self-distillation* prompts a teacher model with access to privileged information, such as a correct solution drawn from the model's own successful rollouts, to produce token-level learning signals. Inspired by this, we propose *verbalized self-distillation* (VSD), a novel method for reweighting negative advantages in GRPO. Rather than conditioning the teacher on a raw correct solution, a *writer model* is given a successful rollout and the failed attempt, and is tasked with identifying the first point of divergence in the attempt and distilling this diagnosis into hindsight instructions. The teacher is then conditioned on these instructions instead of the raw rollout. The per-token divergence between the teacher and the current policy is computed over the policy's failed rollout, and subsequently used to reweight the advantages. Across four mathematical reasoning benchmarks, VSD consistently outperforms GRPO and other self-distillation baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.