Sensitivity-Weighted Reconstruction Mitigates Priority Dilution in LLM Pruning
Abstract
One-shot pruning compresses large language models by matching sparse module outputs to dense counterparts. Existing reconstruction objectives assign unit weight to every calibration token occurrence, despite differences in how local reconstruction errors affect the global loss. We define occurrence sensitivity as the squared norm of the loss gradient at the module output. Across 11 models, the most sensitive 10% of occurrences account for a median of 63% of the sensitivity mass. This concentration exposes Uniform reconstruction to priority dilution: abundant low-sensitivity occurrences can reduce the reconstruction priority of more sensitive ones. We propose Sensitivity-Weighted Reconstruction (SWR), which converts raw occurrence sensitivities into normalized, clipped reconstruction weights and inserts them into existing solvers. With fixed pruning masks, adding the least sensitive half at unit weight raises perplexity in 8 of 11 models; replacing those unit weights with SWR weights reduces this degradation in 10 of 11. Across five pruning methods and 11 models at 70% unstructured sparsity, SWR lowers perplexity in 51 of 55 configurations, with a 3.0% median relative reduction, and improves zero-shot accuracy in 51 of 55 configurations. Further analysis finds that high-sensitivity occurrences are more strongly associated with the token being processed for K/V and generally with the token being predicted next for Q/O/FFN.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.