acceptodds
Under review as a conference paper at ICLR 2027

Prune What the Teacher Can Repair: Distillation-Aware Pruning for LLMs

Abstract

This paper proposes a novel Distillation-Aware Pruning (DAP) framework for addressing the teacher-student compatibility issue in large language models (LLMs). Unlike existing approaches to teacher-student mismatch, which typically assume a predefined student and subsequently adapt the distillation objective to bridge the resulting gap, we revisit the student construction problem through pruning, which provides a natural mechanism for preserving parameters and representations from its parent teacher. Motivated by this intuition, DAP therefore prunes for post-distillation quality rather than immediate post-pruning quality by preserving what the intended distillation process cannot recover and removing what it can repair. Across settings on Qwen3-8B and DeepSeek-R1-Distill-Llama-8B, multiple tasks, and both supervised and on-policy distillation, DAP consistently produces stronger post-distillation students than both independently constructed and conventionally pruned alternatives under matched parameter and training budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.