LSMXQ: Layerwise Salience-Aware Mixed-Precision Quantization for Efficient LLM Inference
Abstract
Post-training quantization (PTQ) is an effective approach for reducing the memory footprint of large language models (LLMs), enabling larger models to fit within limited GPU memory and improving inference efficiency under memory-bound workloads. While recent sub-4-bit quantization methods can achieve high accuracy compared to FP16 baselines, they do not necessarily leverage modern hardware-supported FP4 Tensor Core architectures, limiting their ability to fully utilize the computational throughput of modern GPUs, particularly at larger batch sizes. In this work, we introduce LSMXQ, a hardware-aware mixed-precision post-training quantization approach that allocates low-bit representations under an average bit budget while minimizing layer-wise propagated quantization error. LSMXQ is designed for coalesced memory access and efficient NVFP4 Tensor Core execution, supporting both memory-bound low-batch inference and high-throughput large-batch workloads. Across a broad evaluation of LLMs and zero-shot benchmarks, LSMXQ retains up to 98.5% of FP16 accuracy at 2-bit precision, while matching FP16 accuracy at 2.5- and 3-bit precision. For inference efficiency, LSMXQ achieves up to 3.71× the throughput of FP16 in batch-1 inference, while delivering substantial throughput improvements at larger batch sizes in vLLM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.