acceptodds
Under review as a conference paper at ICLR 2027

Vocabulary-Incomplete Unsupervised Vision-Language Model Adaptation

Abstract

In certain real-world scenarios, obtaining labeled data is excessively expensive, motivating increasing interest in adapting models without annotations. Vision-language models (VLMs) provide pseudo-labels for unlabeled images through image-text similarity, enabling adaptation without manual labeling. The previous unsupervised vision-language model adaption methods assume access to an adaptation vocabulary containing all class names. However, a complete vocabulary is difficult to construct in real-world scenarios, as unlabeled data may contain classes that are overlooked or fall outside the current task scope. Thus, we propose **V**ocabulary-**I**ncomplete **U**nsupervised **V**ision-language model **A**daptation (VI-UVA), where only a subset of class names are available during adaptation. The vocabulary incompleteness induces an adaptation bias toward provided classes, which degrades omitted-class recognition. To address this issue, we propose **L**abel **S**pace **R**econstruction (LSR), which reconstructs a latent label space from the unlabeled data and incorporates it into prompt tuning, thereby extending adaptation beyond the incomplete vocabulary and mitigating the adaptation bias. Furthermore, we derive an omitted-class semantic discrepancy bound for LSR and identify the key factors affecting the bound. Extensive experiments across eight benchmarks demonstrate that LSR outperforms existing state-of-the-art methods. **Code is available in the supplementary materials.**

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.