acceptodds
Under review as a conference paper at ICLR 2027

On the Generalization Limitation of On-Policy Distillation for Post-RL Consolidation

Abstract

Multi-teacher on-policy distillation (MOPD) extends on-policy distillation (OPD) to consolidate capabilities from multiple teachers, but near-perfect in-distribution (IND) distillation can mask incomplete capability consolidation beyond the supervised distribution. A preliminary synthetic reasoning study shows incomplete transfer of teacher generalization. We then experiment on real-world extrapolative and compositional generalization across mathematical reasoning, coding, agent, and long-context capabilities using Qwen3 models. Across extrapolative settings, OPD and MOPD reach 96.04-102.73% of their teachers’ mean IND scores, yet their mean out-of-distribution (OOD) scores remain 2.45-16.25% below the base. Across three compositional settings, MOPD improves mean IND performance by 0.48-22.46% over the base but remains 18.19-24.34% below the best teachers on composed tasks. We analyse teacher capability and data coverage as two potential bottlenecks: jointly easing them eliminates OOD forgetting in our experiments and shows potential for generalization beyond the teacher. Moreover, relative to their RL baselines, OPD and MOPD reduce OOD forgetting on average. Geometry analysis finds more compact OPD and MOPD updates on average than those of their RL baselines, along with greater shared update structure under MOPD. Based on the analysis, we propose geometry-based patching and base-preserving replay to improve OOD performance while retaining most IND performance. Our findings identify a blind spot in capability consolidation via MOPD and motivate distillation algorithms that explicitly support OOD retention and generalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.