acceptodds
Under review as a conference paper at ICLR 2027

CRePE: Convolution-aware Relative Importance in Post-training Pruning with Efficient Search

Abstract

Post-Training Pruning (PTP) reduces the memory and computational costs of Large Language Models (LLMs) by removing weights without retraining. RIA, a state-of-the-art scoring method, normalizes each weight by its row and column sums, but uses only 1D cross-shaped information and weights the two directions equally. We propose CRePE (Convolution-aware Relative Importance in Post-training Pruning with Efficient Search), which adds a 2D local neighborhood term and adaptive per-term coefficients to Relative Importance scoring. Since searching these coefficients by perplexity (PPL)-based hill climbing takes about 11 hours on LLaMA-2-7B, we further propose PHO (Proxy-based Hyperparameter Optimization), which uses the Gini coefficient of pruned importance scores as a surrogate objective and reduces the search to about 20 minutes, a 30 speedup with nearly the same PPL; the resulting coefficients transfer from LLaMA-2-7B to LLaMA-2-13B without re-tuning. Across LLaMA-1/2/3 under 2:4 and 4:8 sparsity and LLaMA-1/2 under unstructured sparsity, CRePE consistently improves over scoring-only baselines, with margins that widen under tighter structural constraints or higher sparsity. Under 2:4 sparsity, CRePE lowers PPL by 0.43 on average relative to RIA. CRePE also combines with channel permutation, non-uniform sparsity allocation, quantization, and re-pruning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.