QUENCH: Prune Iteratively and Quantise Once for Training-Free LLM Compression
Abstract
Large language models require substantial memory to store their weights. Pruning removes selected weights and quantisation represents them with fewer bits, but both can reduce accuracy. A small low-rank correction can recover some of the lost accuracy without fine-tuning. We introduce QUENCH, a training-free pipeline that designs pruning and quantisation around such a correction. We target the linear layers of the transformer architecture. QUENCH alternates pruning and activation-aware correction, enabling the sparse weights to respond to the previous correction. It then quantises the surviving weights once with elastic-bucket quantisation (EBQ), which learns representative weight values (codebooks) for each output channel, and refits the correction. On LLaMA-3-8B with 2:4 sparsity and 3-bit values, QUENCH reaches WikiText2 perplexity 12.01 and 34.22% ARC-Challenge accuracy, against 12.60 and 31.22% reported for EoRA with 2:4 sparsity and 4-bit weights. We also weigh the improvement from activation weighting against low-rank correction. It improves standalone codebooks by 12.29 points, but provides no consistent additional benefit after a rank-128 EoRA correction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.