Generative Vision-Language Models Can Discriminate: Unlocking Zero-Shot Classification Potential
Abstract
Contrastive Vision-Language Models (VLMs), such as CLIP, are naturally suited to zero-shot classification through image–text similarity matching, whereas generative VLMs provide an autoregressive decoding interface that is less direct for closed-set classification. Existing protocols based on open-ended generation or candidate-constrained decoding can be sensitive to label wording, output format, and prompt design. This motivates us to examine whether part of the observed performance gap arises from the inference interface rather than representational capacity alone. In this paper, we propose a zero-shot framework that unlocks the discriminative potential of generative VLMs. By inducing a contrastive-compatible latent space directly from the model's internal representations, we establish a similarity-based matching mechanism that mimics the contrastive formulation, while preserving the rich semantics of the generative paradigm. Furthermore, we identify a phenomenon where the VLM's internal linguistic priors accurately recognize visual concepts but diverge from canonical label taxonomies. To mitigate this, we introduce semantic equivalence rectification, which aligns diverse linguistic manifestations to ground-truth categories. Experiments across multiple benchmarks show consistent improvements over generative prompting baselines. Further analyses identify taxonomy–class alignment as an important, but not exclusive, source of SER's gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.