Gamma-Aware Block Rotations for MXFP4 LLM Quantization
Abstract
Rotation-based post-training quantization reduces outliers by inserting orthogonal transforms and fusing them into adjacent weights. For the residual-stream rotation R1, however, the learned scaling parameter of RMSNorm does not commute with R1 in general, so existing offline methods first fold it into the weights of subsequent linear layers. This operation can amplify magnitude disparities within quantization blocks and degrade the quantized model. We propose Gamma-aware Block Rotation (GaBR). GaBR first permutes channels to group similar values within quantization blocks. It then decomposes the scaling parameters into a block-constant retained in RMSNorm and a residual fused into weights, and alternately optimizes and a block-diagonal R1. With a block-diagonal rotation, retaining one scale per 32 channels in RMSNorm remains exactly equivalent in full precision. Extensive experiments on seven Llama and Qwen models under MXFP4 W4A4 quantization show that GaBR consistently improves upon existing rotation-based PTQ methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.