acceptodds
Under review as a conference paper at ICLR 2027

PRUNE-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning

Abstract

On-policy distillation (OPD) provides dense teacher feedback on student-generated responses, but its rollout budget does not account for how well the student and teacher match along those responses. We identify this unused mismatch information as an opportunity to accelerate OPD: long, weakly compatible suffixes can be curtailed while retaining the supervision needed for learning. We introduce PruneOPD, which uses top-k overlap to measure local compatibility, reweight rewards after cumulative mismatches, and adaptively truncate training rollouts. Within configured length bounds, the algorithm determines and updates the truncation length from compatibility feedback, without a manually specified cutoff or length schedule for each pair. Across four student–teacher pairs, Prune-OPD reduces reported training time by 37.6%–68.0% while broadly preserving accuracy on AMC, AIME, and HMMT. In the high-compatibility DeepSeek-7B / Skywork7B pair, Prune-OPD retains long supervision and stays close to OPD accuracy, avoiding the larger losses from fixed-4K truncation. Prune-OPD thus saves training computation where compatible supervision is short, while retaining longer rollouts where they are needed to preserve performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.