MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models
Abstract
Mixed-precision quantization improves post-training quantization by assigning higher precision to sensitive layers, but existing methods typically optimize bit allocation for a fixed memory budget. On edge devices, the memory available to an LLM can vary at runtime and is unknown during offline calibration. Adaptive quantization addresses this setting by calibrating once and deriving an allocation for the available budget at deployment. However, accurately calibrating before the deployment-time allocation is known is challenging because a layer's distortion depends on the bitwidths of its upstream layers. We show that this dependence substantially changes layer scores and the resulting allocations. We propose MixQuant, a quantizer-agnostic adaptive quantization framework that marginalizes over information unknown during offline calibration while exploiting the fixed quantizer and allocator. MixQuant estimates each layer's expected distortion over quantized upstream configurations, regularizes low-bit assignments that introduce disproportionate distortion, and calibrates quantizer parameters on plans produced by the deployment-time allocator. At deployment, a greedy solver produces an allocation for any available budget in time. Across three LLMs, two base quantizers, multiple memory budgets, and diverse downstream tasks, MixQuant consistently improves perplexity and average task accuracy. On Llama-3.2-3B at , it improves average accuracy over the strongest baseline by and percentage points with AWQ and GPTQ, respectively. Under tight memory constraints, competing methods incur – the degradation from FP16 of MixQuant across the evaluated task categories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.