acceptodds
Under review as a conference paper at ICLR 2027

Effects of Pretraining Concept Frequency and Semantic Similarity among Classes on CLIP Generalized Zero-Shot Accuracy

Abstract

Foundation vision-language models (VLMs), such as CLIP, exhibit impressive generalized zero-shot recognition capabilities. However, their performance is shown to be affected by imbalances in their pretraining data. Understanding the factors that affect downstream zero-shot recognition accuracy under such pretraining imbalance is essential. Recent work has shown that across many downstream datasets, the zero-shot accuracy of individual concepts follows a log-linear trend with concept frequency in the pretraining dataset, and that achieving linear improvements in accuracy requires exponentially more training examples. In this paper, we (i) demonstrate that semantic similarity among target classes is another strong predictor of zero‑shot accuracy, (ii) analyze how this semantic effect interacts with concept‑frequency trends across different dataset sizes, model capacities and pre‑training scales, showing that the frequency‑accuracy link diminishes for larger pretraining corpora while the semantic‑similarity signal does not collapse, and (iii) propose a graph‑based zero-shot accuracy‑estimation framework for a given list of class names that jointly takes (a) concept frequency estimates and (b) a pair‑wise semantic‑similarity matrix as input; allowing for a more flexible and realistic setup than the fixed 1000-way classification. Overall, our results provide a more comprehensive understanding of how both semantic complexity and class cardinality influence generalized zero-shot accuracy of vision-language models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.