Train Where the Model Is Uncertain: Off-Policy Traces for Self-Distillation
Abstract
On-policy self-distillation (OPSD) has emerged as an effective approach for improving large language models. Motivated by the success of on-policy data in reinforcement learning from verifiable rewards (RLVR) and external-teacher distillation, prior work has largely assumed that self-distillation should likewise train on rollouts from the current student policy, leaving off-policy and offline alternatives underexplored. Yet offline training is often more practical. Privileged information such as human feedback may be available only for a fixed corpus of previously collected trajectories and impossible to obtain on live rollouts from the current student policy. Moreover, maintaining a live rollout loop adds substantial computational cost. We study self-distillation when rollouts are generated by policies other than the current student. Surprisingly, offline rollouts outperform on-policy rollouts. We trace this effect to entropy collapse: self-distillation progressively reduces the student's predictive entropy, causing its own rollouts to contain fewer high-entropy tokens—the positions where the privileged-context teacher can provide useful corrections. Offline traces preserve more of these high-entropy positions because they were generated independently of the student's evolving policy. Building on this observation, we introduce two approaches that control rollout staleness to preserve opportunities for the teacher to correct the student. In EMA OPSD, rollouts are generated by an exponential moving average of the student, providing a stale but continually updated rollout policy. In offline OPSD, rollouts come from a fixed corpus generated by other models before training, eliminating rollout generation during training entirely. We additionally introduce high-entropy token masking to focus distillation on uncertain positions. Across mathematical reasoning benchmarks and multi-step Knights and Knaves problems, both approaches substantially outperform standard online OPSD.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.