HyperQuant: A Near-Entropy-Optimal Lattice Quantization Pipeline for Large Language and Diffusion Models
Abstract
Data-free post-training quantization has converged on rotation plus a fixed-rate codebook: whiten a weight tile toward Gaussianity, then round it to an MSE-optimal grid. That last step spends bits per index however unevenly the indices are used. We introduce HyperQuant (Hadamard, optimallY Packing, Entropy Rice-coding), which replaces that codebook with a lattice and a variable-length code in one calibration-free pipeline for both the linear layers and the KV cache of large language and diffusion transformers. A per-tile Walsh–Hadamard rotation whitens each tile, an optimal low-dimensional lattice (, , , or ) quantizes it, a lossless strip discards the bits lattice membership already fixes, and a Rice code carries the remainder to within bps of that lattice's rate–distortion ideal. Decoding is algebraic, so dimension costs no codebook memory. For the KV cache we add subtractive dither, which under an idealized sampler leaves every cached inner product unbiased per vector, and measurably so in practice, where rotation alone is unbiased only on average. Table1 collects operating points. On Llama-3.1-8B weights the default leads HIGGS at – bps, by PPL at over seeds. Neither OCTOPUS nor TurboQuant released usable code, so we reproduce both, validate each against published numbers, and run all three codecs in one harness under one itemized rate accounting over seeds. \method leads both at equal or lower effective rate, apart from a tie with both at bits and a loss to TurboQuant at bits on C4 under QJL; the leads hold at over shared seeds. The same pipeline quantizes the 19B-parameter LTX-2 video DiT with no visible artifacts. On one H100 at batch it cuts resident weight memory and the KV-cache encoding rate (computed rather than measured on a filled cache), though the decode kernel stays slower than bare bf16 GEMV.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.