Off-Diagonal Hessian Sensitivity for Post-Training Quantization of Large Language Models
Abstract
Mixed-precision post-training quantization (MP-PTQ) is a promising approach for improving resource efficiency under constrained computational and memory budgets through effective bit allocation. MP-PTQ methods based on independently optimized layer-wise or module-wise surrogate objectives are computationally efficient, but primarily rely on local quantization errors and therefore do not explicitly account for cross-layer interactions under the task loss. In contrast, search-based methods treat mixed-precision allocation as a black-box optimization problem, leading to costly exploration of a large search space. To address these limitations, we propose Off-Diagonal Energy Sensitivity (ODES), which captures interactions within and across projections through the off-diagonal structure of the task-loss Hessian. Starting from a second-order expansion of the task loss, we separate its off-diagonal interaction term. Then, we derive a sufficient condition under which ODES-guided precision reallocation reduces the second moment of the off-diagonal task-loss contribution. To estimate ODES at LLM scale without materializing the Hessian, we develop a compressed probe-centered (CPC) estimator. CPC reduced estimator state memory by approximately \(126\times\) while producing ODES estimates that closely matched those obtained with the direct Welford estimator. Using the estimated ODES ranking, we reduce mixed-precision bit allocation to a one-dimensional search over the number of reallocated projection pairs and evaluate these candidates using HQQ as a lightweight quantization proxy. This reduced the total allocation time by 36–50% relative to search-based methods while maintaining comparable performance in downstream accuracy and perplexity. ODES-guided allocations also improved over uniform allocation with RTN, HQQ, and GPTQ, indicating that the ODES ranking transfers across different quantizers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.