DFR: Direct Feature Refit for Restoring Depth-Pruned Language Models
Abstract
Depth pruning reduces the inference cost of large language models, but removing Transformer blocks can substantially degrade model quality. Existing training-free methods often address this loss by fitting corrections on the low-dimensional FFN output. This makes fitting substantially cheaper, but can limit recovery quality when useful feature directions fall outside the subspace exposed by the pretrained down projection. We introduce **Direct Feature Refit (DFR)**, which instead fits the pruning-induced correction directly from the surviving intermediate FFN features and merges it into the existing down projection, requiring no fine-tuning or additional inference module. We further derive the globally optimal rank- solution to characterize how the useful correction is distributed across recovery modes. At rank 512, DFR retains 86.39–86.79% of the regularized recovery gain across the three main models, while recovering 94.06–96.47% of the full DFR NLL improvement, showing that most of the language-model improvement is captured by a moderate number of recovery modes. Across two pruning strategies and pruning ratios from 25% to 50%, DFR consistently improves post-pruning model quality on these models, reducing perplexity by up to 53.6% under aggressive pruning relative to the strongest competing baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.