acceptodds
Under review as a conference paper at ICLR 2027

LSMXQ: Layerwise Salience-Aware Mixed-Precision Quantization for Efficient LLM Inference

Abstract

Post-training quantization (PTQ) is an effective approach for reducing the memory footprint of large language models (LLMs), enabling larger models to fit within limited GPU memory and improving inference efficiency under memory-bound workloads. While recent sub-4-bit quantization methods can achieve high accuracy compared to FP16 baselines, they do not necessarily leverage modern hardware-supported FP4 Tensor Core architectures, limiting their ability to fully utilize the computational throughput of modern GPUs, particularly at larger batch sizes. In this work, we introduce LSMXQ, a hardware-aware mixed-precision post-training quantization approach that allocates low-bit representations under an average bit budget while minimizing layer-wise propagated quantization error. LSMXQ is designed for coalesced memory access and efficient NVFP4 Tensor Core execution, supporting both memory-bound low-batch inference and high-throughput large-batch workloads. Across a broad evaluation of LLMs and zero-shot benchmarks, LSMXQ retains up to 98.5% of FP16 accuracy at 2-bit precision, while matching FP16 accuracy at 2.5- and 3-bit precision. For inference efficiency, LSMXQ achieves up to 3.71× the throughput of FP16 in batch-1 inference, while delivering substantial throughput improvements at larger batch sizes in vLLM.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.