Vocabulary-Incomplete Unsupervised Vision-Language Model Adaptation
Abstract
In certain real-world scenarios, obtaining labeled data is excessively expensive, motivating increasing interest in adapting models without annotations. Vision-language models (VLMs) provide pseudo-labels for unlabeled images through image-text similarity, enabling adaptation without manual labeling. The previous unsupervised vision-language model adaption methods assume access to an adaptation vocabulary containing all class names. However, a complete vocabulary is difficult to construct in real-world scenarios, as unlabeled data may contain classes that are overlooked or fall outside the current task scope. Thus, we propose **V**ocabulary-**I**ncomplete **U**nsupervised **V**ision-language model **A**daptation (VI-UVA), where only a subset of class names are available during adaptation. The vocabulary incompleteness induces an adaptation bias toward provided classes, which degrades omitted-class recognition. To address this issue, we propose **L**abel **S**pace **R**econstruction (LSR), which reconstructs a latent label space from the unlabeled data and incorporates it into prompt tuning, thereby extending adaptation beyond the incomplete vocabulary and mitigating the adaptation bias. Furthermore, we derive an omitted-class semantic discrepancy bound for LSR and identify the key factors affecting the bound. Extensive experiments across eight benchmarks demonstrate that LSR outperforms existing state-of-the-art methods. **Code is available in the supplementary materials.**
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.