DTDD: Divergence-Triggered Dynamic Distillation for Reliable On-Policy Supervision
Abstract
On-policy distillation (OPD) provides dense token-level supervision on student rollouts, but this supervision can become unreliable when the student and teacher diverge, either because of a large initial policy gap or compounding errors in long-horizon reasoning and tool use. We introduce Divergence-Triggered Dynamic Distillation (DTDD), an adaptive OPD framework that routes supervision between standard OPD and teacher-generated recovery using segment-level student–teacher divergence. When divergence triggers an intervention during reasoning, the student prefix continues to receive standard OPD supervision. The teacher then generates a recovery suffix, which the student learns through behavior cloning, an online SFT-style loss. We evaluate DTDD in three regimes: long-horizon reasoning, distillation from a base initialization without offline SFT, and multi-turn agents. We also compare alternative trigger designs and teacher-intervention rates. DTDD reduces the post-warmup peak gradient norm by 8.6× and, from a base initialization, reaches 21.3% held-out math accuracy in 2.6× fewer optimizer updates than OPD; both methods finish near 22%. It improves best-observed performance by 4.8 and 4.1 percentage points on ALFWorld and agentic code, respectively. Ablations show that segment-level triggers outperform token-level triggers. Behavior cloning is also more effective than the standard OPD objective for training teacher-generated recoveries. Across 23 experimental arms, takeover tends to help when the teacher substantially outperforms the OPD-trained student, provided that the teacher can recover from student-visited states, the student can learn those recoveries, and teacher intervention is neither excessive nor insufficient.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.