acceptodds
Under review as a conference paper at ICLR 2027

PACE: Precision Allocation with Cost Estimation for Mixed-Precision Quantization

Abstract

Mixed-precision quantization reduces model storage by assigning different bitwidths under a fixed budget. Its effectiveness depends on accurately estimating which precision decisions best preserve model quality. Existing methods often rely on indirect sensitivity measures that do not reflect the actual output damage caused by different quantization candidates. We introduce PACE (Precision Allocation with Cost Estimation), a second-order mixed-precision quantization method grounded in a stationary output-fidelity objective. PACE efficiently predicts the output damage of concrete precision candidates and globally allocates the available bit budget according to their estimated benefits. Experiments validate the consistency between predicted and measured quantization damage and compare PACE across three 7–8B model families, two perplexity benchmarks, six zero-shot tasks, and average 6-bit and 3-bit settings. Under the average 6-bit setting, PACE ranks first or second across all nine evaluations, with WikiText2 perplexity increases of at most 0.095, while remaining robust under the more challenging average 3-bit setting. On Qwen3-8B, PACE is approximately faster than SFMP, the strongest comparable-quality baseline. Its packed representations provide approximately compression under W6 and more than compression under W3.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.