acceptodds
Under review as a conference paper at ICLR 2027

Byte-Constrained Layer-Wise Quantization with Cost-Normalized Sensitivity

Abstract

Mixed-precision post-training quantization allocates storage across tensors of different sizes and sensitivities, but nominal bit width omits format metadata. We instantiate a byte-constrained multi-choice knapsack formulation with measured sensitivity and runtime-supported formats, then allocate precision using quality gain per byte. Across five models from three families, the resulting GGUF artifacts lower perplexity relative to six production llama.cpp presets at 25 of 30 matchedsize points. Paired block-bootstrap intervals resolve 23 improvements and one regression. Controlled ablations separate storage accounting from ranking quality. Uncorrected nominal budgets overshoot file-size targets by 10.6% to 17.9%, but nominal and byte pricing produce similar perplexity after size matching. In contrast, raw-sensitivity ranking is worse at 29 of 30 matched-size points, by up to 6.36 perplexity. Across six downstream tasks, Benjamini–Hochberg correction at 5% identifies 21 gains and three regressions among 180 comparisons, with 19 gains on models of at most 3B parameters. A same-model untying control attributes most shared-embedding sensitivity to the language-model head. These results distinguish accurate artifact sizing from effective precision allocation, while exposing the limits of a separable sensitivity objective.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.