AngularQ: Geometry-Guided Gradients for Low-Bit Large Language Models
Abstract
Training large language models (LLMs) with one- to three-bit weights and activations offers major gains in efficiency, but preserving accuracy under extreme quantization remains a central challenge. The difficulty lies not only in representing discrete values but also in learning through them, since hard quantizers eliminate the derivatives required for gradient-based training. We introduce AngularQ, a mathematical framework for quantization-aware training that leverages the averaging effect of random embeddings to obtain a smooth population formulation of discrete low-bit computation. This formulation yields geometry-aware gradient surrogates that retain information from the quantized forward pass while providing an unbiased surrogate for the corresponding backward gradient. We characterize how different embedding constructions affect approximation accuracy and gradient quality, using these insights to guide a unified, computationally efficient approach to binary and multibit training. Beyond providing a comprehensive theoretical foundation for learning through quantization, this framework translates into substantial empirical gains in extreme-low-bit language modeling. Across language models from 45M to 1.7B parameters on FineWeb- Edu, AngularQ delivers its largest gains when both weights and activations are binary: 55.3% and 49.4% lower perplexity than matched-update QuEST at 0.6B and 1.7B, respectively. Two- and three-bit configurations retain substantial gains of 7.7–9.3%. Three paired training seeds at 140M consistently reproduce the binary advantage. At equal code-storage budgets, multibit quantization also outperforms binary projection expansion across the evaluated scales. These results establish angular geometry as a principled and effective foundation for learning extreme-low-bit language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.