acceptodds
Under review as a conference paper at ICLR 2027

DepthQuant: Depth Error Propagation for Mixed-Precision Bit Allocation

Abstract

Mixed-precision quantization assigns a bit width to each unit of a large language model under a fixed bit budget, but most current methods decide the allocation from single-layer information and fall well short of full precision. This may be partly because single-layer methods miss information from deeper layers: similar quantization errors can produce different responses at the model output. To address this, we propose DepthQuant, a mixed-precision quantization method that uses depth error propagation as the signal that drives the bit allocation. By measuring the response each precision change induces one layer deeper and how much it changes the model output, DepthQuant spends the bit budget where more precision reduces that measured change the most. Built on this principle, DepthQuant introduces a cohesive multi-scale design: short-depth response scoring ranks precision changes by the response they induce one layer deeper, full-depth allocation then adjusts the resulting allocation by the error's effect on the model output, and a final step sets each row's range from information in its own layer. Experiments show that DepthQuant improves on the compared baselines in most evaluated settings and in every reported average-2-bit setting, reducing perplexity by up to 40.6% at 2.5-bit average precision relative to the best-performing baseline and raising zero-shot accuracy by up to 11.6% at 2-bit average precision, where these methods degrade the most.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.