Beyond Local Error: Functional Reconstruction for Vision-Language Model Quantization
Abstract
Post-training quantization (PTQ) aims to reduce the cost of vision-language models while retaining their capabilities. Reconstruction-based methods pursue this goal by matching quantized intermediate outputs to their full-precision counterparts, with modality-aware approaches further balancing visual and textual errors. Yet an error’s impact depends on what happens next: subsequent operations can respond differently to perturbations of similar magnitude. A smaller local error therefore need not better preserve the computation that follows. We introduce Modality- Native Functional Equalization (MNFE), a PTQ method based on multimodal functional reconstruction. MNFE evaluates quantization errors at module-specific states formed by downstream operations and balances their contributions across visual evidence, textual context, and answer information. A structured search refines channel scales through bounded group-wise adjustments. Extensive ex- periments on Qwen2.5-VL-7B and InternVL2.5-8B show that MNFE achieves the highest nine-benchmark average among evaluated quantized methods under W3A16, W4A8, and W4A6. Notably, under W4A6, MNFE retains 94.8% of the full-precision average score on InternVL2.5-8B. Our code will be made publicly available upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.