Constrained Refinement of Layerwise Sparsity Control for LLM Pruning
Abstract
Post-training pruning of large language models (LLMs) requires deciding which weights to remove within each layer and how to distribute a fixed sparsity budget across layers. At high sparsity, different layerwise allocations under the same global budget can yield substantially different perplexity (PPL) after pruning. Many existing methods derive layerwise sparsity ratios from importance or sensitivity scores or from predefined allocation rules, providing useful starting points. However, these initial allocations may not account for the layerwise reconstruction response of the selected pruning backend at the assigned sparsity levels. This motivates Plug-and-Prune, a constrained method that refines a uniform allocation or one produced by an existing allocator. At the input allocation, it measures layerwise reconstruction responses under the selected backend and uses them in a regularized update of the keep ratios. The update penalizes deviations from the input allocation and enforces the global budget and per-layer bounds, while leaving within-layer mask construction to the backend. We evaluate Plug-and-Prune with WANDA and SparseGPT on LLaMA, Qwen3, and Phi-4 models, primarily at 70% unstructured sparsity, supplemented by sparsity-sweep and structured-pruning experiments. The results show that our method generally lowers PPL relative to the input allocation, with gains tending to be larger at higher sparsity. Overall, response-guided budget redistribution is a practical, backend-agnostic complement to existing inner-layer criteria, especially at high sparsity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.