acceptodds
Under review as a conference paper at ICLR 2027

The Undetected Damage of Quantization on Retrieval and How to Fix It

Abstract

Quantization rounds weights to a lower bit-width, and it is often judged by accuracy loss. This criterion is incomplete. The same quantized model that preserves classification accuracy changes 20-46% of top-1 retrieval results, and the aggregate ranking metrics reflect only part of that change. The top-1 result is guaranteed to survive quantization only when the gap between the two highest scores, logits in classification and query-document similarities in retrieval, exceeds twice the largest rounding error. In classification, the loss function pushes the correct class apart from every other, ensuring this gap, but in retrieval nothing separates the top-1 item from the second. This gap can be measured without labels. Before deployment, it predicts which models will break, and at deployment time it tells, per query, whether the quantized answer still matches the one full precision would have returned. This holds for most classification inputs and a minority of retrieval queries. That gap motivates a different fix in each setting. In retrieval, spending bit-width on the layers most sensitive to it recovers three-fifths to three-quarters of an extra bit's benefit for half its cost. In classification, routing the few low-gap inputs to full precision recovers most of the lost accuracy at a fraction of the cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.