acceptodds
Under review as a conference paper at ICLR 2027

E8-QAT: Efficient 2-Bit Quantization-Aware Training with Implicit Lattice Codebooks

Abstract

Low-bit quantization reduces the memory footprint of large language models, but preserving accuracy with efficient training remains challenging. Scalar quantization encodes weights independently, while explicit-codebook vector quantization introduces costly search during training. We propose E8-QAT, a 2-bit quantization-aware training framework that uses the lattice as an implicit eight-dimensional codebook. It jointly quantizes weight groups through arithmetic projection without storing or searching an explicit codebook during the training forward pass. A 2-bit core encoding and sparse extension preserve the selected lattice points exactly, adding less than 0.1% to core index storage. On Qwen3.6-35B-A3B, E8-QAT matches the performance of the full-precision model in terms of the average score across the four tasks. E8-QAT improves over the state-of-the-art baseline by 2.03% on IFEval and by 1.51% on LiveCodeBench, under the same data, schedule, and bit width. The 2-bit representation reduces peak memory by 80% and supports a larger batch. At each representation's largest feasible batch, E8-QAT reaches the decoding throughput of BF16.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.