acceptodds
Under review as a conference paper at ICLR 2027

CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search

Abstract

The best two-bit weight quantizers for large language models, such as QTIP and Proteus, share a recipe. Each weight matrix is first rotated by a random orthogonal transform so that its entries look Gaussian, and this rotation has to be undone at every decoding step. The rotated weights are then encoded with a trellis or lattice code under a Euclidean search, and the layer Hessian is used only in the error-feedback step between coding blocks. We show that this recipe leaves part of the Hessian unused. Error feedback turns the loss into a weighted sum of per-coordinate rounding errors, and the weights, the diagonal of the Hessian's LDL factorization, are computed by existing quantizers but never read. We put these weights into the search's distance, the Viterbi branch metric, so that the search follows the curvature within each coding block. This also explains what the rotation is for. The rotation removes this within-block variation, so after it the weight has little left to exploit. Weighting in the native basis and rotating are therefore substitutes. We confirm this on three models: the weighted native search matches the baselines' rotation, a full-dimension randomized Hadamard, to within about one point of downstream accuracy, and weighting after the rotation gains little. Around this search we build CurveTQ, a trellis codec with no rotation. The rotation also flattened the amplitude and the marginal shape of the weights, which we handle directly with a factored scale field and a closed-form quantile table. In either basis, error feedback passes each block's residual on to the next, so we store a start state per coding block that lets the trellis adapt to it. At two bits CurveTQ is 1-3 points higher in mean downstream accuracy than QTIP and Proteus on three 4-8B Instruct models, even after both are given our start state, which by itself lifts either baseline by 1-3 points. It also leads on a 35B mixture of experts, to our knowledge the first trellis-coded result on such a model. With no rotation to undo, our decoder is the fastest of the three at all tested batch sizes and bit widths.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.