A-OPSD: Activation-Space On-Policy Self-Distillation for Large Language Models
Abstract
On-policy self-distillation (OPSD) improves reasoning models by using a privileged teacher —the same model conditioned on both a problem and its reference solution— to supervise a student that observes only the problem. Existing OPSD methods transfer this privileged signal through dense token-level distribution matching over the student's own rollouts. While effective, this supervision acts only after the internal reasoning of the student and teacher has been projected through the language modeling head, and its cost scales with the full vocabulary. We ask whether the same signal can be transferred directly in the model's hidden-state space, and study a family of activation-space objectives for OPSD—from plain hidden-state matching to metrics derived from the local geometry of the OPSD loss. We find that even naively matching the teacher and student's hidden-states can match token-level OPSD, while a Gauss–Newton (Fisher) pullback of the OPSD KL through the LM head can outperform it. Across Qwen3 and OLMo-3-7B models with sizes ranging from 1.7B to 8B, A-OPSD (Activation-Space On-Policy Self-Distillation) matches or exceeds token-level OPSD on mathematical reasoning benchmarks (MATH-500, AIME24, HMMT25) while requiring less supervision memory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.