GapSQ: Closing the Pruning-Induced Gap for Joint Sparsification and Quantization of LLMs
Abstract
Post-training sparsification and quantization can be combined to compress large language models, but pruning changes the weight distribution seen by the quantizer in a systematic way. By removing many near-zero weights, it creates a pruning-induced gap between the retained negative and positive branches, leading to under-utilization of the quantization range and available levels. Motivated by this observation, we introduce GapSQ, a training-free method for aggressive joint compression. We primarily study 50% sparsity with 3- or 4-bit quantization, where GapSQ narrows the pruning-induced gap through group-wise inward shifting before quantization. To avoid storing per-weight branch information, GapSQ enforces quantized-space separability, allowing the shift to be reversed using only lightweight group-level metadata. Across Llama and Qwen models, multiple sparsity patterns, and several pruning–quantization backbones, GapSQ consistently improves joint-compression baselines, with larger gains observed in more aggressive low-bit regimes. For example, on Llama 3-8B with 3-bit quantization and 50% sparsity, GapSQ improves average zero-shot accuracy from 57.2% to 62.5% while reducing WikiText-2 perplexity from 30.52 to 11.72 relative to SparseGPT+GPTQ. These results show that explicitly adapting post-pruning weight distributions is an effective way to improve low-bit joint compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.