Two for the Price of One: Curvature-Aware Pruning Preserves Influence Functions Too
Abstract
Post-training model pruning enables efficient deployment, but conventional methods do not explicitly preserve internal-gradient alignment between sparse and dense models. We study novel connections between downstream performance of compressed models and the ability to retain gradient information via influence-function preservation. Central to this connection is a curvature-aware objective that accounts for both layer activations and backpropagated gradients. Despite its promise, this objective remains underexplored in pruning due to a lack of efficient solvers. We introduce , a novel ADMM-based optimizer supporting both pruning and sparse-plus-low-rank decompositions of pre-trained weights. Compared to LLM-Surgeon which targets the same objective, achieves over an order-of-magnitude speedup. At % pruning of Llama-3.2-1B, reduces the C4 perplexity gap to the dense model by %, improves compressed-to-dense per-sample gradient cosine similarity on WikiText-2 from to , and runs over faster than LLM-Surgeon on substantially weaker hardware. By preserving influence functions, supports effective data selection. On Llama-3.2-3B, fine-tuning a fresh dense model on the % of data selected by achieves % accuracy on BIG-Bench Hard, compared with % using ALPS and % using the dense-mode selection. Finally, integrating with LoGRA, a recent gradient-projection strategy, on Qwen3-4B with MLP layers pruned using 2:4 sparsity and a rank-64 correction retains a cosine alignment with dense LoGRA-projected C4 gradients.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.