From Gradient Surgery to Bottleneck Regularization: A Unified Framework for Combining Verifiable-Reward RL and On-Policy Self-Distillation
Abstract
Reinforcement learning with verifiable rewards (RLVR) gives reasoning models a clean signal about whether a rollout is correct, but says little about which tokens made it correct. On-policy self-distillation (OPSD) offers the opposite trade-off. An answer-conditioned copy of the model can provide dense token-level supervision, but that privileged teacher can also exploit information the student will not have at test time. Recent work further shows that teacher-student likelihood shifts need not track token value. We study how to combine these signals without letting privileged supervision override the verifier. CIBO-CRISP, our primary method, uses token-level REINFORCE with a bounded, polarity-aware, sign-preserving credit multiplier, a softly gated teacher-student divergence penalty, and an exponential-moving-average stability anchor. CRISP and VACS-CRISP are simpler intermediate constructions that expose the design choices leading to the final objective. On Qwen2.5-Math-1.5B, CIBO-CRISP averages 62.7% Pass@1 on MATH-500 across two seeds, compared with 44.6% for GRPO, while remaining close on GSM8K (74.1% versus 72.4%). In matched seed-0 ablations, the full method reaches 63.6% on MATH-500 versus 60.8% with fixed beta and 60.2% with the privileged-divergence term removed. The theoretical claims are deliberately narrower than the practical algorithm; this is because we give a standard fixed-surrogate nonconvex-SGD result and an algebraic leakage-bookkeeping identity, but do not treat the implemented KL term as a mutual-information estimator or teacher likelihood shifts as causal token values.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.