COTQuant: Cross-Modal Orthogonal Transform for Large Vision-Language Model Quantization
Abstract
Post-training quantization (PTQ) is essential for efficient deployment of vision-language models (VLMs), yet low-bit quantization remains challenging due to heterogeneous multimodal activation outliers. We propose COTQuant, a PTQ framework with two complementary components. Cross-Modal Transform (CMT) jointly redistributes text and vision activation outliers through hierarchical orthogonal transformations, while Quantization-Aware Group Search (QAS) reorganizes weight channels to minimize group-wise quantization error. Both transformations leave the model output unchanged and introduce only a small amount of additional inference-time overhead. Across three VLMs and five multimodal benchmarks, COTQuant consistently outperforms existing PTQ methods while retaining 98.03%–98.73% of the corresponding BF16 performance under W4A4 quantization. Notably, we demonstrate effective 4-bit PTQ on Qwen3.5-122B, which is the first 4-bit PTQ attempt on over 100B-scale VLM.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.