acceptodds
Under review as a conference paper at ICLR 2027

Schedule the Optimizer, Not the Teacher: Diagnosing and Mitigating Long-Horizon Collapse in On-Policy Self-Distillation

Abstract

On-policy self-distillation (OPSD) lets a single language model teach itself. A frozen copy of the initial policy, conditioned on privileged ground-truth solutions, provides dense per-token supervision on the student's own rollouts. This delivers gains competitive with group relative policy optimization (GRPO) at a fraction of the sampling cost. We ask what the short budget hides. Training every frozen-teacher OPSD variant past 100 steps, we find a consistent trajectory: performance peaks, declines monotonically, and by step 300 falls below the untrained base model, at 1.7B, 4B, and 8B scales. The obvious suspect is the frozen teacher going stale, but the obvious remedy backfires. Refreshing the teacher to track the student collapses the model to a three-benchmark mean of 5.7 (8.9% on AIME24, vs. base 54.7%); its outputs degenerate into fluent but empty self-talk. Gentler anchor motion is no escape. An exponential moving average (EMA) teacher matches the best early score of the frozen-teacher family, yet ends the horizon below the frozen teacher's own endpoint at every rate we test (32.5–32.9 vs. 35.1 across EMA 0.99–0.9995). Only dose decay absorbs the penalty. The frozen teacher is not the bottleneck but the anchor. The collapse is instead objective drift under a constant learning rate, controlled primarily by the cumulative learning-rate dose. The one-sided clip that stabilizes early training explains the drift's visible signature, namely a training loss that descends below zero. It does not explain the drift itself. A horizon-aware cosine decay removes most of the collapse at all three scales, ending at or above the untrained base at 1.7B and 8B. Alongside the diagnosis, we make three supporting contributions. A difficulty-aware sample weighting brings small multi-seed gains and does not alter the long-horizon outcome. A numerically hardened hybrid of self-distillation and reinforcement learning with verifiable rewards trains stably where the naive form suffers a 10^14 gradient explosion that poisons the optimizer state. A defect analysis of the official implementation clarifies which reported numbers are affected. Our results reframe frozen-teacher self-distillation as a short-horizon method: schedule the optimizer, not the teacher. Our code is available at the anonymous repository: https://anonymous.4open.science/r/Schedule-the-Optimizer-Not-the-Teacher-ICLR-2027-DCB6.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.