TORQUE: Optimizing What (not) to Quantize Before and After Rotation
Abstract
Random rotations are an effective preprocessing step for quantization: regardless of the input, they make coordinate distributions approximately Gaussian, enabling the use of precomputed optimal codebooks. We introduce TrunQuant, a framework that improves on previous work using random rotations by jointly optimizing how many and which coordinates to preserve at high precision before and after rotation, under a fixed overall expected bit budget. Intuitively, before rotation, preserving large input coordinates at high precision can reduce overall error by preventing the rotation from spreading their values across many coordinates. Likewise, after rotation, preserving a small fraction of the largest-magnitude coordinates at high precision allows the remaining values to be quantized more accurately using codebooks optimized offline for the resulting truncated Gaussian distribution. We demonstrate TrunQuant's effectiveness across various biased and unbiased quantizers, including both scalar and vector methods, and show that it improves worst-case accuracy and applications such as KV-cache compression and approximate nearest-neighbor search.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.