CoCaQ: Coupling of Compression and Quantization for Flexible LLM Deployment
Abstract
The practical deployment of Large Language Models (LLMs), especially on edge devices, is heavily constrained by their massive model sizes and prohibitive computation costs, which are generally addressed by either model compression or quantization. In practice, these two kinds of techniques can be applied concurrently for flexible LLM deployment, especially under extremely low memory budgets, which, however, are typically optimized separately and so lead to suboptimal performance. Particularly, we observe that the joint of model compression and quantization is challenged by the effect of compressed structures on quantization and the ability of quantization to account for distribution shifts induced by compression. To address the above issues, we propose a new Coupling of Compression and Quantization () method to perform model compression and quantization in a coupling manner, aiming to achieve effective and flexible LLM deployment by unifying the two techniques. To this end, our CoCaQ is developed with two key components, including a quantization-aware model compression (QaMC) and an activation stabilization for quantization (ASQ). First, our QaMC constructs quantization-friendly low-rank representations based on activation statistics and optimizes the layer-wise compressed structures with a quantization-aware loss, aiming to identify the robust substructures of LLM for the subsequent quantization. Furthermore, our ASQ introduces an efficient orthogonal rotation regularization for low-rank weight matrices, enabling low-bit quantization to better tolerate activation amplification induced by the preceding model compression. Based on the proposed QaMC and ASQ, our CoCaQ couples model compression and quantization by fully considering the mutual influence between the two, resulting in an effective yet flexible solution for LLM deployment. Experiments are conducted across LLaMA and Qwen models, and the results show that our CoCaQ achieves better memory-accuracy trade-offs than state-of-the-art compression-only and quantization-only methods, while clearly outperforming the counterparts with joint of compression and quantization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.