acceptodds
Under review as a conference paper at ICLR 2027

Representation Matters in 2-Bit QAT: Finite Barnes–Wall Representations for Large Language Models

Abstract

Quantization-aware training (QAT) is widely used to recover the quality of large language models (LLMs) at ultra-low precision, yet most existing methods focus on improving how a given low-bit representation is optimized. We study a complementary question: what representation should QAT optimize in the first place? We introduce BW-QAT, a representation-first framework for 2-bit weight-only QAT based on a finite 16-dimensional Barnes–Wall representation. BW-QAT encodes every 16 weights with a fixed 32-bit payload, calibrates the gain alphabet directly against the finite reconstruction space, and initializes trainable proxies from compensated pre-quantization targets. This initialization preserves the exact initial hard-quantized model while retaining the residual information needed for effective discrete updates. The resulting representation can be trained with hard quantized forward passes and a standard straight-through estimator. Across Qwen and Llama models, BW-QAT consistently outperforms matched lattice, scalar-QAT, and vector-QAT baselines in the low-data regime. On Qwen3-1.7B, it improves six-task average accuracy from 52.17% to 58.30% over a matched LC-QAT adaptation and reduces WikiText-2 perplexity from 18.90 to 14.02 across three training seeds. These results show that representation quality is a first-order factor in data-limited 2-bit QAT.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.