acceptodds
Under review as a conference paper at ICLR 2027

Improve MXFP4 Quantization via Inter-block Permutation and Micro-block Givens Rotation

Abstract

Large language models (LLMs) have achieved remarkable performance across diverse tasks, but their growing computational and memory costs pose substantial challenges for efficient deployment. Post-training quantization (PTQ) has emerged as a practical solution, while recent hardware support for microscaling formats such as MXFP4 further enables efficient low-bit inference. However, MXFP4 remains highly sensitive to massive activation outliers due to its power-of-two E8M0 shared scales and low-precision E2M1 values. Existing block-wise Hadamard transformations mitigate outliers by uniformly spreading their energy across an entire quantization block, but such magnitude reduction is not necessarily optimal for MXFP4. We observe that MXFP4 contains a set of favorable, exactly or nearly representable magnitudes, termed **lossless quantization points**, and quantization error is strongly determined by the distance of transformed values to these points. Based on this observation, we propose a novel transform framework for MXFP4. First, we propose *Inter-block Permutation* to group massive outliers with low-importance channels, reducing their interference with important features under block-wise shared scaling. Second, we propose *Micro-block Givens Rotation* to adaptively control the number of participating channels and redistribute outlier energy toward favorable quantization points, instead of uniformly dispersing it across the entire block. Experiments on the LLaMA2/3 and Qwen3 model families demonstrate that our method better adapts to the quantization characteristics of MXFP4 than the Hadamard transform and achieves superior quantization performance across multiple LLM benchmarks, providing new insights into the design of quantization methods for next-generation microscaling low-precision formats.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.