Structured Category-wise Selective On-Policy Distillation
Abstract
On-policy distillation (OPD) aligns teacher and student predictions on student-generated trajectories, but conventional token-level objectives entangle semantic-category choices with preferences among alternative realizations within the same category. We propose Structured Category-wise Selective On-Policy Distillation (SCS-OPD), which separates category-mass alignment from intra-category supervision. Category-level alignment drives the main rollout update; an optimizer-matched finite-step selective gating mechanism then compares category-plus-residual and category-only updates, retaining intra-category residuals only when the gate indicates a lower category-level loss. Across OPD experiments covering multiple methods, the Qwen and DeepSeek model families, and multiple random seeds, SCS-OPD achieves the best performance on most mathematical benchmarks. In the settings where it leads, SCS-OPD improves over methods including Reverse-KL OPD, EOPD, and TSD by 0.3–3.4 percentage points, while maintaining stable performance on out-of-domain benchmarks. Mechanistic analysis shows that, after semantic aggregation, intra-category supervision is sparse, highly concentrated, and heterogeneous. Selective gating over fine-grained realizations therefore provides substantial optimization benefits for OPD.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.