SQuAT: Spherical Lattice Vector Quantization-Aware Training
Abstract
Vector quantization (VQ), in particular lattice-based methods, offer strong rate-distortion trade-offs for low-bit large language models (LLMs), yet their behavior under quantization-aware training (VQ‑QAT) remains poorly understood. We scale lattice VQ‑QAT to billion-parameter LLMs and multi-billion-token training regimes, and introduce a shape–gain VQ‑QAT approach, SQuAT, that decomposes weights into direction and magnitude: normalized directions are quantized using an E8-derived spherical codebook, while gains are quantized independently with low-bit scalar quantizers. We analyze the interaction of shape-gain parameterizations and straight-through estimators (STE) in the context of VQ optimization, and show that factored STEs induce tangent-space updates for directions, yielding an implicit form of spherical (Riemannian) optimization. To enable scalability, we develop a distributed exact nearest-neighbor assignment scheme, in which workers quantize disjoint weight blocks and communicate only discrete indices, reducing VQ-QAT training time by 12.8x and allowing us to scale it to 10B tokens. Across instruction-tuned LLaMA-3.2-1B/3B, Qwen3-4B, and base Qwen3-1.7B, SQuAT consistently outperforms standard VQ‑QAT STE under identical quantizers and initialization, and achieves SOTA VQ-QAT accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.