Effective-Supervision-Aware On-Policy Distillation
Abstract
On-policy distillation (OPD) provides token-level teacher supervision on student-generated trajectories. Yet these trajectories differ in answer correctness and in their agreement with teacher predictions, raising a practical question: how should training updates be allocated among them? We study effective-supervision-aware allocation in OPD: using outcome and teacher-based signals to prioritize trajectories for training without changing the distillation loss. As a concrete implementation, we introduce the Reliable Trajectory Buffer (RTB). RTB uses answer correctness and teacher-assigned likelihood to select current trajectories for replacement, teacher–student predictive compatibility to admit previously encountered mistakes to a replay buffer, and rescoring after replacement to refresh distillation signals. Across six mathematical reasoning benchmarks, with Qwen3-8B as teacher, RTB improves the aggregate score from 25.66 to 28.01 for Qwen3-1.7B-Base and from 35.68 to 38.03 for Qwen3-4B-Base. It also reaches higher scores earlier when measured by newly sampled student trajectories. Ablations support the roles of selective replacement and compatibility-based admission.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.