acceptodds
Under review as a conference paper at ICLR 2027

LoRScore: Low-Rank Refinement of Activation-Aware Scores for Learnable Large Language Model Pruning

Abstract

Learnable LLM pruning adapts sparsity patterns to the language-modeling objective, helping preserve model quality under compression. However, existing approaches often attach trainable variables to individual weights or candidate mask patterns, introducing substantial optimization overhead. We propose LoRScore, a parameter-efficient framework that combines Wanda's activation-aware importance scores with a learnable low-rank residual while keeping pretrained weights frozen. The key idea is to constrain only the learned score correction, preserving fine-grained prior information while enabling coordinated, loss-driven adjustments to pruning decisions. By separating score learning from mask projection, LoRScore supports both unstructured and semi-structured sparsity within a shared parameterization and introduces no additional parameters at inference. Experiments across multiple LLM families and sparsity settings demonstrate competitive language-modeling and zero-shot performance with substantially fewer trainable parameters. Theoretical analysis connects the surrogate score gradients to first-order sensitivity and characterizes the induced low-rank updates, providing an interpretation of how the compact parameterization refines pruning decisions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.