acceptodds
Under review as a conference paper at ICLR 2027

LoPRo: Enhancing Low-Rank Quantization via Permuted Block-Wise Rotation

Abstract

Post-training quantization (PTQ) enables effective model compression with minimal accuracy loss. While most weight-only PTQ methods target the challenging sub-3-bit regime and often require fine-tuning, we propose LoPRo—a fine-tuning-free algorithm that improves residual matrix quantization via block-wise permutation and Walsh-Hadamard transforms to align columns of similar importance. We also introduce a mixed-precision, fast low-rank decomposition based on rank-1 sketching (R1SVD) to reduce quantization costs. Experiments show LoPRo outperforms existing fine-tuning-free PTQ methods at 2- and 3-bit (with its vector variant LoPRo delivering the best 2-bit results), matching fine-tuned accuracy while achieving up to 5 quantization speedup over GPTVQ. On Mixtral-8x7B, it finishes quantization in under 2.5 hours, reducing perplexity by 0.75 over MoEQuant++ (3.17 vs. 3.00 avg. bits) and boosting BoolQ accuracy by 5.4 points. LoPRo also achieves higher accuracy than other low-rank methods at much lower ranks, with low latency overhead (10.3% at batch 1, lower for larger batches).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.