acceptodds
Under review as a conference paper at ICLR 2027

Residual-Aware Sparsification and Pruning: Protecting Principal Subspaces During Sparsification or Pruning

Abstract

Linear layers dominate LLM parameter count and memory traffic, making them a major computational bottleneck. Existing structured pruning and activation-sparsity methods reduce these costs by removing or skipping channels according to saliency or activation-based criteria. However, such methods can discard dominant singular directions of the original projection, making performance sensitive to the scoring rule or sparsification criterion. We propose Residual-Aware Sparsification and Pruning (RASP), a training-free drop-in **wrapper** for pruning and sparsity methods. For each targeted weight \(\mW\), RASP decomposes \(\mW = \mL_r + \mR_r\), where \(\mL_r\) captures the top-\(r\) singular directions and \(\mR_r\) is the residual. RASP then confines pruning or sparsification to \(\mR_r\), preserving the dominant projection directions while bounding approximation error by the residual norm, rather than the full weight norm. At 50% sparsity, RASP consistently improves generation scores – by up to +7.33 points over GRIFFIN, +7.93 over Wanda, +2.10 over TEAL. Interestingly, RASP surpasses dense inference on several Gemma 2 9B IT generation benchmarks. Finally, RASP is able to achieve 40%-46% speed-ups compared to a dense model. Our code is publicly available at https://anonymous.4open.science/r/RASP-EB5E.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.