acceptodds
Under review as a conference paper at ICLR 2027

GLAS: Gradient-Guided Layer-wise Sparsity Allocation for LLM Pruning

Abstract

Large language models (LLMs) have achieved remarkable success across diverse tasks, yet their colossal size renders deployment expensive in both memory and inference. Pruning offers a direct remedy, but prevailing techniques apply a uniform sparsity ratio across Transformer layers, disregarding their varying importance and incurring severe performance degradation at high sparsity. Recent efforts therefore shift toward non-uniform layer-wise pruning. However, they rank layers by empirically designed proxy scores and overlook the dependence of layer importance on the pruning metric in use. To address this, we propose GLAS (Gradient-Guided Layer-wise Sparsity Allocation), which casts sparsity allocation as a constrained optimization problem. GLAS- allocates sparsity guided by metric-dependent gradients in a single pass, while GLAS- further refines the allocation iteratively. Experiments across mainstream LLM families show that GLAS surpasses state-of-the-art baselines in both perplexity and zero-shot accuracy, and transfers to semi-structured and structured pruning as well as to vision transformers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.