acceptodds
Under review as a conference paper at ICLR 2027

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

Abstract

Structured pruning is a hardware-friendly way to compress LLMs, but preserving recognition scores does not ensure useful free-form generation. On-Policy Distillation (OPD) offers a recovery signal by having the frozen pre-compression model supervise the student's own response prefixes. In the pruning regime we study, however, long rollouts frequently develop repetitive suffixes. The teacher can also assign these suffixes high probability, so low distillation loss can coexist with poor generation. We propose ShortOPD, an OPD rollout schedule guided by terminal repetition and clean truncation. Repetition permits reducing the next budget toward an estimated prefix length, whereas non-repetitive responses that reach the cap permit growth. The controller reuses existing tokens and teacher scores, requires no additional model forward pass, and preserves the current batch's full-token distillation update. Its short-to-long phase follows an initial contraction from the maximum budget. Across four Qwen3 model settings with 25% of layers removed, ShortOPD attains 82.3–84.6% of the teacher's eight-task aggregate generation score, outperforming Vanilla OPD by 11.4–18.2 points. Under an 8192-token ceiling, it also uses 14.7–30.2% fewer rollout tokens while scoring 13.95–16.21 points higher. We hope this recipe helps move structured pruning beyond marginal gains on perplexity and multiple-choice benchmarks, a step closer to deployment-ready generation quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.