Vodka: Vector quantization-aware training via Differentiable Kernel Approximation
Abstract
The low bits quantization of Large Language Models (LLMs) is essential for lowering deployment costs and enabling inference with limited resources. As vector quantization (VQ) offers higher expressivity than scalar quantization (SQ) and quantization-aware training (QAT) demonstrates a higher upper-bound than post-training quantization (PTQ), combining VQ with QAT shows better potential for low-bit quantization. However, in face of the optimization and efficiency problems, current VQ-QAT methods either constrain the expressivity, or require extensive search, which motivates us to explore the tradeoff between expressivity and efficiency. Therefore, we propose Vodka, a novel VQ-QAT algorithm via differentiable kernel approximation. We identify the key expressivity gap as the mapping function: current VQ-QAT restricts this map to be affine, whereas VQ allows it to be arbitrary. To bridge the gap, Vodka constructs the Codebook Spectrum with kernel by ordering the curvature, with two endpoint functions provably the affine and arbitrary, which stays search-free and admits an exact codebook gradient at fixed indices, serving as implicit codebook to approach the VQ controllably. In contrast to prior work with a fixed initialized codebook, we optimize the implicit codebook jointly with the weights, so that both adapt to each other. We also propose Shakernel, a fused GPU kernel that implements Vodka without heavy intermediate tensors, reducing training cost and memory. Empirical evaluations on Qwen3.6-35B-A3B and Llama3.1-8B on 7 benchmarks demonstrate that Vodka outperforms state-of-the-art (SOTA) low-bit methods while maintaining efficiency. Notably, Vodka improves over the strongest baseline by 1.31% on IFEval and 0.51% on BFCL, demonstrating the excellent performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.