SafeQ: Safety Alignment Preserving Post-Training Quantization for Language Models
Abstract
With the growing adoption of Language Models (LMs), Post-Training Quantization (PTQ) has become a popular standard for efficient LM deployment, especially to resource-constrained environments like personal and edge devices. Most PTQ methods only consider to preserve natural language understanding (understanding) capabilities, without taking safety capabilities into consideration, that enables LMs to refuse harmful, unethical, privacy-breaching, or illegal user queries. This inadvertently causes complete or partial safety alignment degradation in quantized LMs. To address this, we propose SafeQ, a learnable PTQ method that preserves safety alignment by taking both understanding and safety capabilities of LMs into consideration. Precisely, SafeQ learns quantization parameters by minimizing a block-wise weighted-sum loss over understanding and safety capabilities for both weight-only quantization and weight-activation quantization. We performed extensive evaluations for assessing understanding capabilities, and safety and general utility capabilities using different pre-trained and instruction-tuned LMs respectively. Our evaluations demonstrate best or near-best performance by SafeQ in pre-trained LMs for W4A16 (i.e., 4-bit weight and 16-bit activation), and W4A4 quantization modes, and in instruction-tuned LMs for W4A16 quantization mode. In few instances, SafeQ improved safety capabilities of quantized LMs slightly beyond full-precision LMs. Our results evidently show that SafeQ resolved trade-off between safety and general utility capabilities to a great extent in instruction-tuned LMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.