CSQuant: Cross-Block Smoothing for Accurate and Efficient Low-Bit LLM Quantization
Abstract
Low-bit quantization is essential for efficient large language model (LLM) inference, but aggressive quantization remains vulnerable to activation outliers that amplify quantization error. Orthogonally rotating activations alleviates quantization error but incurs additional online computation, particularly in adaptive methods that use intermediate permutations, which prevent multiple rotations from being combined and introduce explicit data movement. We propose CSQuant, an accurate and efficient post-training quantization framework comprising two key modules: Cross-Block Smoothing and LiteBoA. Cross-Block Smoothing replaces intermediate permutations with a normalized cross-block Hadamard transform, enabling cross-block outlier redistribution while allowing two adaptive rotations to be merged into a single online transformation. This separable structure further enables CUDA kernel fusion with adjacent operations, reducing intermediate tensor materialization and global-memory traffic. LiteBoA selectively applies attention-aware second-order compensation to value projection matrices and employs head-shared Hessians to reduce calibration cost while preserving quantization accuracy. Experiments on LLaMA-1/2/3 demonstrate improved quantization accuracy across W4A4KV4 and weight-only settings, with particularly strong performance at 2-bit weights. CSQuant achieves a 2.25 prefill speedup over FP16 under W4A4KV4 quantization with a sequence length of 2048 and a batch size of 4, while our fusion strategy further accelerates existing Hadamard-based rotation methods. CSQuant also substantially reduces peak GPU memory during post-training quantization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.