acceptodds
Under review as a conference paper at ICLR 2027

Vista: Accurate CLIP Model Inversion Framework with Semantic Targets and Appearance Coverage for Data-Free Quantization

Abstract

How can we quantize a CLIP model without access to real data while preserving its generalization capability? Quantization reduces the memory and computational cost of large models, but most methods require real data for calibration. Data-free quantization (DFQ) eliminates the need for real data, yet most existing methods synthesize images conditioned on predefined class labels as inversion targets. In contrast, CLIP models support open-vocabulary recognition without a fixed label space. Therefore, DFQ for CLIP models requires semantically diverse text prompts to extract broad semantic knowledge encoded in the model. A recent method manually constructs these targets, covering a narrow semantic range and degrading the generalization of the quantized model. In this paper, we propose Vista, an accurate CLIP model inversion framework for DFQ. Vista broadens the semantic coverage of inversion targets by systematically sampling visual prompts from the lexical hierarchy of WordNet. Vista captures semantic relationships among prompts by replacing one-hot supervision with identity-preserving soft targets derived from pairwise prompt similarities. Vista expands the visual coverage of synthetic images by decorrelating the appearance components that the target prompts do not explain, while enforcing view consistency to preserve semantic alignment. Experimental results show that Vista improves the average accuracy of quantized CLIP models by up to 9.45%p over existing DFQ methods and better preserves the generalization capability of CLIP models across various CLIP backbones, bit-widths, and downstream datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.