Density-Matched Power-of-Two Quantization: Allocating Logarithmic Levels for Multiplier-Free LLM Inference
Abstract
Large language model (LLM) inference spends most of its compute in the multiply-accumulate (MAC) operations of linear projections. Post-training quantization (PTQ) reduces this cost with low-bit integer weights and activations, leaving a multiplier in every MAC. Power-of-Two (PoT) weight quantization replaces that multiplier with a shifter but spaces its levels a factor of two apart, degrading accuracy at low bit-widths. Finer fixed-ratio grids narrow the gap but remain suboptimal, since they spend the same number of levels per doubling of magnitude, wherever the weights lie. We introduce **Density-Matched Power-of-Two (DMPoT)** quantization, which formulates logarithmic level allocation as a continuous point-density optimization problem. We derive the density that minimizes quantization mean squared error (MSE) and use it to guide the construction of a multi-piece family of fractional PoT grids whose deployable members support folded shift-add hardware. Three results follow. (i) The derived density concentrates quantization resolution on the upper flank of the weight distribution while leaving the bottom of the range sparse. (ii) The optimal grids at successive bit-widths nest as sets of levels into one chain, and two bit-width laws fix the operating regime at four to five bits: a one-adder-per-weight bound and a width-minimality result at INT-matched MSE. (iii) Every deployed 4-bit level is an integer shift of or of , so a folded datapath computes the single addition once per activation, shared by all output channels. Across seven models and two configurations, DMPoT matches the uniform integer baseline in all fourteen cells, never more than perplexity above it, with mean gaps of and . The synthesized MAC array with 4-bit weights and activations occupies the area of its integer counterpart.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.