Reopening Reinforcement Learning Headroom in Language Models with Iterative Reset On-Policy Distillation
Abstract
Reinforcement learning with verifiable rewards can substantially improve language-model reasoning, yet continued optimization often enters a low-gain regime. We introduce **Iterative Reset On-Policy Distillation and Reinforcement Learning** (**IR-OPD-RL**), which repeatedly resets a student to a fixed supervised fine-tuning checkpoint, transfers the current teacher's behavior through reward-extrapolated G-OPD, and launches a new GRPO stage from the resulting student. This design transfers behavioral knowledge without inheriting the teacher's parameters or optimizer history. Across Qwen3-1.7B and Qwen3-4B on six mathematical-reasoning and code-generation configurations, IR-OPD-RL improves domain-average accuracy over the initial RL teachers by 3.30-8.23 points. On Qwen3-4B, IR-OPD-RL further exceeds ExOPD, the strongest tested single-stage baseline, by 5.42 points on Math-57K and 2.45 points on Code-25K. Under matched compute on Qwen3-1.7B, IR-OPD-RL outperforms continued RL by up to 2.35 points on Math-57K and 3.70 points on Code-25K. Order and initialization ablations further favor distillation before RL and resetting the student from the fixed SFT anchor. These results show that immediate checkpoint accuracy need not predict subsequent RL improvement and that reset on-policy distillation can renew optimization headroom.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.