REVQ: Rotation-Enhanced Product Vector Quantization with Low-Rank Residual Compensation
Abstract
Joint low-bit quantization of weight and activation could substantially reduce the storage and computational costs of Transformer-based architectures and convolutional neural networks (CNNs). However, existing methods cannot simultaneously minimize the layer-wise local quantization errors caused by joint misalignment between weight and activation distributions and address accumulation of errors across layers. In this paper, we formulate the cross-layer and local-layer errors for joint weight and activation quantization, and consequently, propose a rotation-based two-stage framework named REVQ for low-bit quantization to simultaneously address the errors. Specifically, we develop learnable groupwise rotation transformations parametrized by the Lie algebra of the special orthogonal group SO(g) to learn the coordinate system of parameter (weight and activation) space for achieving product vector quantization and reducing local-layer quantization error. Moreover, we design cross-layer low-rank residual compensation to bridge rotation-based quantization in adjacent linear layers with learned low-rank compensation via Haar wavelets for minimizing cross-layer quantization error. Extensive experiments demonstrate that REVQ achieves state-of-the-art performance in various tasks of image classification on ImageNet and CIFAR-100, object detection and instance segmentation on COCO, and language modeling on Wikipedia using a wide range of Transformer-based architectures and CNNs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.