Stability-Aware Structured Pruning of Large Language Models with Dynamic Search Spaces
Abstract
Structured pruning reduces the memory and computational costs of large language models (LLMs) by removing complete attention heads and multilayer perceptron (MLP) neurons. However, importance estimates derived from limited calibration data can vary across samples, and fixed search boundaries restrict adaptation to optimization feedback. We propose DSASP, a structured pruning framework that combines cross-sample stability assessment with dynamic search spaces. DSASP augments first-order Taylor and empirical-Fisher importance estimates with two complementary statistics: Importance Variance, which measures fluctuations in score magnitude, and Cross-sample Ranking Consistency, which measures stability in relative ordering. A Tree-structured Parzen Estimator optimizes the four scoring weights using proxy perplexity within an adaptive search region. Improvements shift and contract this region, whereas sustained stagnation expands it. Across LLaMA-7B and Vicuna-7B at pruning ratios of 20% and 40%, DSASP achieves the lowest perplexity among the evaluated pruning methods on both WikiText-2 and PTB and the highest mean zero-shot accuracy in all four settings. Ablation studies support the contribution of both the combined stability signals and the dynamic search mechanism. These results establish cross-sample stability and adaptive configuration search as complementary tools for preserving model performance under structured pruning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.