OffQ: Taming Structured Outliers in LLM Quantization by Offsetting
Abstract
Low-bit quantization is now a standard approach for accelerating large language model (LLM) inference, as it substantially reduces memory usage and compute cost. However, structural activation outliers remain a major obstacle to effective quantization. In this paper, we introduce OffQ, a post-training quantization method that mitigates this issue through a novel offsetting mechanism. Specifically, OffQ first identifies a low-dimensional outlier subspace in the activations using a tailored top-1 PCA, and then concentrates high-magnitude activations into a single channel by rotation. It then absorbs this concentrated outlier energy into a shared offset, reducing the variance of the activations and making them more amenable to uniform-grid, uniform-precision quantization. This strategy enables accurate W4A4KV4 quantization of LLMs while preserving efficient low-bit execution. Experiments across diverse LLM architectures and benchmarks show that OffQ consistently outperforms state-of-the-art baselines, improving accuracy while maintaining low-bit efficiency. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.