Q-PAT: Quantization-Aware Pruning-Aware Tuning for Low-Bit Large Language Models
Abstract
Structured pruning and low-bit quantization are two key techniques for reducing the deployment cost of large language models. However, structural decisions made during pruning may misalign with the quantization errors introduced later, potentially degrading model performance under low-bit precision. In this work, we propose Q-PAT, a quantization-aware pruning-aware tuning method tailored for low-bit deployment. Building on the Pruning-Aware Tuning (PAT) framework, Q-PAT constructs a channel-wise quantization-sensitivity proxy based on the INT4 fake-quantization error of low-rank weights. The resulting sensitivity signal is incorporated into a Mask-Quant regularization term through a stop-gradient mechanism, thereby constraining the optimization of trainable structural masks at the loss level without directly injecting quantization perturbations into the main forward pass. This design enables mask optimization to account for heterogeneous low-bit quantization sensitivity across channels during task adaptation, better coordinating task adaptation, structural compression, and quantization error. Experimental results show that Q-PAT preserves strong downstream task performance under low-bit quantization and achieves improved INT4 performance on representative Gemma-7B settings. In particular, Q-PAT improves the average score across 14 downstream tasks from 56.69% to 57.38% and increases MMLU accuracy from 24.04% to 33.64%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.