PrismQ: Noise-Aware Resource Allocation for Elastic Post-Training Quantization
Abstract
Elastic post-training quantization (PTQ) jointly calibrates multiple precisions within a single model, enabling flexible precision adaptation under different deployment budgets. However, extending reconstruction from a single precision to multiple precisions introduces two key challenges. First, quantization errors vary across linear layers and precisions, creating a trade-off between cross-precision parameter sharing and precision-specific reconstruction capacity within the reconstruction space. Second, calibration states should reflect precision-dependent error propagation, while preserving complete quantized states for every precision incurs substantial calibration memory overhead. To address these challenges, we propose PrismQ, a noise-aware elastic PTQ framework that adapts both reconstruction space and calibration states according to precision-dependent quantization errors. Specifically, PrismRank estimates the demand for reconstruction capacity at each linear layer and precision from the singular-value characteristics of clipping-induced errors, and PrismLoRA realizes the resulting capacity allocation through shared and precision-specific low-rank components. Moreover, Calibration-only Sparse Token selects quantized token states according to token-wise quantization error under a fixed cache budget and uses them to construct calibration inputs for the next block, allowing these inputs to reflect precision-dependent error propagation. Extensive experiments on language models, vision transformers, and vision-language models show that PrismQ improves quantization performance across multiple precisions, with particularly clear gains at low bit widths.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.