Correlation-Aware Structured Pruning for Large Language Models
Abstract
Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Existing methods often rely on independent unit scores or capture cross-unit dependencies only implicitly, which may misestimate the joint effect of pruning correlated units. To address this, we propose a Correlation-Aware Structured Pruning method. We explicitly decompose the set-level reconstruction error into individual pruning costs and pairwise interactions that capture dependencies in activations and weights, leading to a cardinality-constrained binary quadratic program. Since this binary quadratic program is NP-hard and difficult to solve exactly, we develop a greedy interaction algorithm based on dependency-aware marginal costs to optimize unit selection. Furthermore, we incorporate a gradient-based strategy to achieve adaptive layer-wise sparsity allocation across the entire model. Extensive experiments on mainstream LLMs demonstrate that incorporating correlation information yields competitive accuracy-efficiency trade-offs compared to representative structured pruning baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.