HiCAP-Q: Hierarchical Capability-Aware Mixed-Precision Quantization for Large Language Models
Abstract
Large language model quantization effectively reduces inference costs and memory usage. However, existing methods typically minimize weight or activation reconstruction errors during calibration. These objectives are only weakly aligned with preserving downstream capabilities. As a result, a small reconstruction error does not necessarily imply a small degradation in domain-specific abilities such as mathematical reasoning, code generation, or long-context understanding. We propose **HiCAP-Q**, a capability-aware mixed-precision post-training quantization framework. It preserves model-specific capabilities under a user-specified bit budget. Our key insight is that different capabilities rely on different model parameters. Some layers and linear modules are critical for mathematics, coding, or long-context reasoning. Others can tolerate more aggressive quantization. **HiCAP-Q** exploits this structure by reallocating precision from capability-insensitive parameter groups to capability-critical ones. **HiCAP-Q** consists of three stages. First, it constructs capability-specific probing inputs at low cost and without human annotation. It uses the target model and, optionally, a judge model. An **input score** selects prompts based on question quality, task relevance, and response confidence. Second, the framework estimates capability sensitivity at the Transformer-layer and linear-module levels. A capability loss score measures the degradation caused by quantizing each parameter group. Third, the framework ranks parameter groups by sensitivity and assigns mixed precision under a target bit budget. Experiments on MMLU-Pro, AIME-2026, and GPQA show consistent gains over reconstruction-based baselines. With only 10% more average bits than uniform low-bit quantization, **HiCAP-Q** substantially narrows the gap to the full-precision model. It also preserves strong general-task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.