acceptodds
Under review as a conference paper at ICLR 2027

Compress to Improve: Reclaiming Output Budget from On-Policy Distillation for Stronger Reasoning

Abstract

On-policy distillation (OPD) teaches a student model to reason through teacher feedback on its own generated responses. We find that this process also produces substantial length inflation: across six OPD-family methods evaluated from a common base model on competition mathematics, responses are consistently longer than those produced by supervised fine-tuning (SFT) and reinforcement learning (RL). Under a finite output budget, longer initial solutions leave less room for checking, revision, and further exploration. We introduce concise post-training as an intermediate stage that favors shorter correct solutions and frees output space for further reasoning. Across OPD, ExOPD, and OPSD, it reduces mean output length by 29.6–36.6% while improving both sample accuracy and problem coverage. We then reuse the released budget through prompted continuation, which asks the compressed model to check and revise its answers with fixed parameters, or renewed OPD, which resumes teacher-guided learning. Continuation improves sample accuracy in every route, while renewed distillation increases problem coverage. The strongest final endpoint, ExOPD followed by compression and continuation, raises average@8 from 25.2% to 27.3% and pass@8 from 36.7% to 40.8% relative to the original ExOPD checkpoint. Both complete paths outperform their original distilled models without exceeding their mean output lengths, supporting compression as an intermediate stage for further reasoning and learning gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.