acceptodds
Under review as a conference paper at ICLR 2027

Prune-Quantize-Prune: Towards Better Joint Model Compression by Getting Free Sparsity from Symmetric Quantization

Abstract

The deployment of large language models (LLMs) at scale imposes substantial demands on memory bandwidth and computational throughput, making model compression essential for efficient inference. Among post-training compression techniques, 2:4 semi-structured sparsity and low-bit quantization are particularly attractive, yet effectively combining them remains challenging. Existing approaches typically apply pruning and quantization sequentially, either in a prune-then-quantize or quantize-then-prune manner, without explicitly exploiting the additional sparsity naturally induced by quantization. To address this limitation, we propose a three-stage **prune-quantize-prune** framework tailored to the combination of semi-structured sparsity and symmetric quantization. In the first stage, we introduce a *Quantization-aware Mask Selection* strategy that avoids prematurely enforcing the exact 2:4 constraint and instead reserves room for quantization-induced free sparsity. After applying low-bit symmetric quantization, the resulting additional zeros are explicitly exploited. In the final stage, we re-prune the quantized weights to strictly satisfy the target 2:4 sparsity pattern. To mitigate the reconstruction error introduced by this post-quantization pruning, we further propose a *Grid-constrained Compensation* method, which performs second-order weight compensation directly in the quantized space while preserving the discrete quantization grid. Extensive experimental results on popular LLMs demonstrate that this method outperforms existing state-of-the-art methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.