ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights
Abstract
We introduce ShamAN-Q, a sub-1-bit post-training quantization method that extends NanoQuant (Chong et al., 2026) by replacing its diagonal reconstruction geometry with a dense but tractable curvature metric, using a form and computation paradigm popularized by the Shampoo optimizer (Gupta et al., 2018). For each linear weight, ShamAN-Q fits a Kronecker product to the empirical Fisher Information matrix of a small calibration set by Kullback–Leibler minimization and uses the result to form a Mahalanobis reconstruction loss. The continuous ADMM updates from NanoQuant become solutions to Sylvester equations, while its discrete projection and deployment format remain unchanged. Because the curvature is local to a given set of weights, ShamAN-Q re-measures the input curvature statistic for each layer immediately before layer factorization and periodically refreshes all statistics on the partially quantized model. ShamAN-Q also redistributes the uniform rank from NanoQuant across layers at the same total number of bits. On Qwen3-Base, ShamAN-Q lowers WikiText-2 perplexity at 1 bpw from 27.56 to 22.96 (0.6B), 19.21 to 16.72 (1.7B), and 14.29 to 13.80 (4B) while increasing zero-shot reasoning task performance on the Eleuther LM Evaluation Harness. On 0.6B, ShamAN-Q at 0.8 bpw matches the published perplexity of NanoQuant at 1.0 bpw.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.