acceptodds
Under review as a conference paper at ICLR 2027

Quantization Beyond Matrix Multiplication for Efficient LLM Inference

Abstract

Quantization is widely used to accelerate LLMs inference, but existing approaches primarily optimize matrix multiplication while leaving surrounding operators and data movement in high precision. This limits end-to-end benefits of quantization, particularly for memory-bound operators. In this paper, we proposed an INT8 pipeline that extends quantization beyond matrix multiplication. Our design preserves intermediate tensors in INT8 across both matmul and non-matmul operators, reducing global-memory traffic. We further introduce a hierarchical warp–block reduction strategy that effectively utilize on-chip resources and maintains high GPU occupancy during compute deep learning operators. We evaluate our approach on Qwen3 and Llama-3.2 models across multiple tasks and workloads. Compared with BF16 inference, our approach achieves up to inference speedup and reduces energy consumption by up to 39%, while maintaining comparable accuracy. The benefits remain consistent across GPUs, decoding workloads, and batch sizes, demonstrating that quantization beyond matmul is important for efficient end-to-end LLM inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.