SparRotQuant: Patterned Sparsity-Aware Rotation-Based Quantization for LLM Compression
Abstract
Low-bit quantization and semi-structured N:M sparsity are two hardware-friendly techniques for efficient large language model deployment. However, their direct combination suffers from a structural conflict: rotation-based quantization smooths weight distributions to improve low-bit representations, whereas N:M sparsity relies on strong local contrast within each weight group to determine a stable keep/prune pattern. Consequently, applying N:M pruning after rotation produces unstable sparse masks and significant quality degradation. We propose SparRotQuant, a sparsity-aware rotation-based quantization framework that incorporates the N:M sparsity constraint into calibration, aligning the learned sparse support with the rotated low-bit representation used for deployment. To guide this process, we introduce Probability-aware Balanced Group Separation (PBGS), a group-level criterion that combines retention probability with Hessian-aware sensitivity to encourage each weight group to form a reliable sparse pattern before hard projection. With frozen pretrained weights, SparRotQuant learns differentiable sparse masks during calibration while adapting rotations and quantization parameters to the resulting sparse low-bit representation. Experiments across multiple model families show that SparRotQuant maintains strong model quality under W4A4KV16 and W4A4KV4 while preserving the efficiency benefits of sparse low-bit execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.