Prompting What Matters: Adaptive Calibration for Data-Free Quantization of Large Vision-Language Models
Abstract
Data-Free Quantization (DFQ) provides a promising solution for compressing Large Vision-Language Models (LVLMs) without real calibration data. Existing DFQ methods mainly focus on global statistical matching, visual semantic preservation, or generic contrastive alignment, overlooking the heterogeneous quantization sensitivity of multimodal capabilities and model components. As a result, semantically plausible synthetic samples may fail to activate capability-specific quantization errors that dominate performance degradation. In this work, we reveal that LVLM quantization vulnerability is highly capability-module dependent, raising two fundamental questions: which capabilities and modules are most sensitive to quantization and how can calibration samples be synthesized to recover them? To address these challenges, we propose PromptQ, Prompt-Guided Data-Free Quantization framework that optimizes calibration distribution according to capability-module quantization vulnerability. PromptQ first introduces Sensitivity-Driven Prompt Construction (SDPC), which estimates capability-module sensitivity through behavioral degradation and representation discrepancies under controlled quantization. Based on the estimated sensitivity, SDPC allocates calibration prompts across vulnerable capabilities while maintaining semantic diversity, and generates module-aware prompts that specify the target capability and quantization-sensitive component. Guided by these prompts, Quantization-Sensitive Image Synthesis (QSIS) synthesizes semantically valid calibration samples that expose informative quantization errors by increasing the discrepancy between full-precision and quantized representations under semantic constraints. By coupling sensitivity-aware capability allocation with module-targeted sample synthesis, PromptQ transforms data-free calibration from distribution matching into vulnerability-aware error discovery. Extensive experiments on diverse multimodal benchmarks demonstrate that our PromptQ consistently improves quantization performance without requiring real calibration data. Our code is available in supplementary material package.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.