NS-UDPD: A Finite-Domain Joint Selection Method for Layer-Wise Sparsity Allocation in LLM Pruning
Abstract
Layer-wise sparsity allocation improves unstructured post-training pruning of large language models (LLMs) by distributing a global sparsity budget across Transformer layers. Existing metric-based methods overlook pruning-cost variations across sparsity levels. Search-based and learning-based methods lack global optimality guarantees for discrete allocations, while reconstruction-based methods optimize Euclidean error rather than the language-modeling negative log-likelihood (NLL) objective. We therefore propose **NS-UDPD, a finite-domain joint selection method** that selects one candidate sparsity level per layer while exactly matching the target mean layer sparsity. **Diagonal-Kronecker NLL Sensitivity (DKNS)** combines squared weights, input-activation second moments, and squared output-NLL-gradient energies into per-weight DKNS proxy deletion costs, which candidate deletion masks aggregate into layer-level DKNS state costs. Cost-Only Allocation may concentrate state costs in shallow layers, where pruning perturbations propagate through subsequent layers. **Uniform Depth-Prefix Dominance (UDPD)** therefore bounds the cumulative state cost at every depth prefix using the uniform reference allocation. A three-level lexicographic objective first minimizes the full-depth DKNS state cost and then improves the depth-prefix cost distribution and layer-wise sparsity balance. We evaluate NS-UDPD on five models from the LLaMA-3 and Qwen3.5 series under three fixed local pruning rules and 40%–70% global unstructured sparsity. At 70% sparsity, it reduces WikiText-2 perplexity relative to the best competing allocation on Qwen3.5-2B with SparseGPT by 31.55% and raises six-task mean accuracy from 37.20% to 41.26% on Qwen3.5-4B with Wanda relative to Uniform.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.