acceptodds
Under review as a conference paper at ICLR 2027

MirrorQuant: Learnable Householder Rotations for Low-Bit LLM Quantization

Abstract

Orthogonal rotations have become an effective technique for low-bit LLM quantization by expressing hidden representations in more quantization-friendly bases. However, existing approaches predominantly parameterize learned rotations as dense matrices, incurring quadratic learned state and costly optimization. Their deployment efficiency largely depends on foldability: at non-foldable interfaces, dense rotations must remain on the inference path, with both quadratic storage and computation of the rotation dimension. To overcome these limitations, we propose MirrorQuant, a unified framework for efficient and learnable rotations across Transformer interfaces. MirrorQuant parameterizes learned rotations using short chains of Householder reflections, a construction that preserves exact orthogonality, reduces the learned state to linear in the rotation dimension for fixed chain depth, and enables learnable log-linear online execution without materializing dense rotation matrices. We further introduce projection-guided learning, which complements representation reconstruction by penalizing quantization distortion after the immediate downstream linear projection. Experiments across 23 matched settings spanning model families, scales, and low-bit configurations demonstrate accuracy and efficiency gains. MirrorQuant improves 0-shot accuracy over DartQuant in every setting, with gains up to 1.81%, while its online rotation reduces prefill latency by 29.6% relative to a dense implementation. Source code can be found in the supplementary materials.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.