acceptodds
Under review as a conference paper at ICLR 2027

SoftWater: The Rate-Distortion Geometry of Softmax Quantization

Abstract

Post-training quantization methods typically optimize weighted mean-squared error (WMSE) on linear-layer outputs, yet the final softmax layer is fundamentally different: its relevant distortion is the change in the predictive distribution. We formulate softmax-layer quantization as a rate-distortion problem under KL divergence and show that, to second order, its geometry is class-aware, coupling feature covariance with class-specific softmax curvature. This distinction matters because output heads are often left in high precision even though, in modern language models, they can account for 15–30% of all parameters, and over half of the stored bytes once the body is quantized. We introduce SoftWater, a practical quantizer derived from this KL geometry. A separable approximation reduces the otherwise prohibitive joint class-feature weighting to a single feature-side factorization together with inexpensive class-dependent rescaling, allowing the head to be quantized using statistics collected in one calibration pass. The resulting allocation assigns precision according to the output class frequency discounted by its variance across inputs, which is especially important in highly imbalanced output spaces such as LLM vocabularies. Across five language models from 1B to 32B parameters, SoftWater outperforms the released WaterSIC quantizer, near-optimal under WMSE, at matched head rates on 59 of 60 evaluation points, reducing head-induced KL by 6.5–8.3 at 2 bits despite using none of WaterSIC's additional refinements. Because the class-side statistic comes from calibration data, matching calibration to the deployment domain gives the lowest KL on that domain throughout. On Llama-3.2-1B-Instruct with quantized bodies, where the tied head is also the embedding, a 2-bit head removes 45–60% of stored bytes for a 2.9–3.7% perplexity increase and a 4-bit head is near-lossless. These results show that treating the softmax layer according to its own distortion geometry can make aggressive end-to-end model quantization practical.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.