CAPQuant: Covariance-Aware Partitioning of Activation Channels for 4-bit LLM Quantization
Abstract
Low-bit post-training quantization (PTQ) reduces the memory and latency footprint of large language models (LLMs), but activation outliers can compromise accuracy. Group-wise activation quantization limits their reach by assigning one scale to each fixed-size block of channels, and rotation-based methods flatten the outliers before quantization. Which channels share a scale is then a free choice that these methods leave to the arbitrary channel order. We introduce CAPQuant, a training-free method that selects this order from the activation means and cross-channel covariances of an already-transformed model. Starting from the expected error of asymmetric group-wise quantization, we bound the per-group range by the within-group dispersion and reduce the dispersion to a sum of pairwise channel distances. Across balanced partitions the per-channel variances contribute a constant, so only mean differences and covariances decide the grouping, and sorting channels by their mean is already optimal when they are uncorrelated. CAPQuant minimizes the remaining objective with balanced clustering and pairwise swaps at calibration time, then folds the permutation into the adjacent weights and fuses it into the activation kernels, so inference needs no additional matrix multiplication or activation pass. On top of OSTQuant, converting the per-token activation sites to groups and ordering their channels reduces perplexity by 1.5–9.0% and raises zero-shot accuracy by 0.9–2.3 points across Qwen3-4B, Llama-3.2-1B, Llama-3-8B, and Llama-2-7B, with gains that persist at 13B and 30B parameters and transfer to frozen QuaRot and SpinQuant checkpoints. Ordering alone accounts for up to 2.5% of perplexity beyond contiguous grouping, and on Qwen3-4B its share grows as groups shrink, reaching 2.6% at the MXFP4 group size . The fused permutation adds 3.6–6.9% to prefill latency in a BF16 reference implementation. These results support channel order as an inexpensive calibration-time improvement for group-wise quantization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.