acceptodds
Under review as a conference paper at ICLR 2027

Toward Robust Human-free Image Annotation with Thousands of Candidate Classes

Abstract

Vision-language models (VLMs) can reduce manual image labeling, but annotation becomes difficult when the label space grows, and the target data are imbalanced. Question-answering VLMs are a natural choice, yet their computational cost increases and their annotation accuracy decreases as the candidate vocabulary grows: on ImageNet subsets, accuracy drops from 73.9% at 200 classes to 49.5% at 1,000. Retrieval-based VLMs such as Contrastive Language-Image Pre-training (CLIP) and Sigmoid Loss for Language-Image Pre-training (SigLIP) scale more efficiently. However, standard zero-shot inference still treats each image independently. We therefore ask: *can the unlabeled target pool itself improve annotation without target labels or knowledge of the true class distribution?* Our analysis shows that class names and target images provide complementary information: the class name is a stable semantic *anchor*, while the pool reveals target-specific visual structure and class prevalence. Standard pseudo-label prototyping only partly exploits this signal, because confidence filtering biases the visual representative and uniform balancing fails under class imbalance. We propose *robust prototyping*, a closed-form semantic–visual calibration that assigns images under an estimated class prior, builds a whitened visual head from all assigned images, and fuses it with the semantic anchor. Without target-pool labels, our method improves zero-shot retrieval on all 18 CLIP/SigLIP encoder-dataset pairs, outperforms the label-free prototype baseline on 15 of 18, and achieves the best results among the evaluated training-free methods on the *long-tailed* ImageNet-LT and Places-LT benchmarks, with only seconds of computation after feature encoding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.