acceptodds
Under review as a conference paper at ICLR 2027

OPDARENA: UNDERSTANDING LEARNING AND FORGETTING IN ON-POLICY DISTILLATION

Abstract

On-policy distillation (OPD) is increasingly used for language model post-training, yet how teacher supervision shapes learning and forgetting remains unclear. We present OPDArena, a framework for comparing supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and OPD across five pretrained models and diverse teacher sources, jointly evaluating mathematical learning, general capability retention, and predictive distribution changes. We find that teacher choice strongly shapes retention: RLVR-trained teachers support task learning while preserving prior capabilities, whereas SFT-trained teachers induce greater capability loss. Out-of-domain (OOD) distribution change shows a more consistent positive association with forgetting than in-domain change. We then examine the additional information teachers provide relative to student Base. Matched teacher comparisons connect preferences learned through RLVR to lower relative OOD influence and stronger task learning with limited forgetting; calibrated local directions further support differences in cross-domain effects. Parameter and prediction trajectories show how students acquire teacher preferences through different parameter paths, with task gains and forgetting inherited at different rates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.