Beyond Pseudo-Labels: Vision-Language Guidance for Generalized Category Discovery
Abstract
Generalized category discovery (GCD) aims to recognise known categories while discovering and clustering novel ones in a pool of unlabelled data. Previous methods rely on the visual similarity between samples to form clusters; however, this fails in applications where the distinctions between categories are fine-grained. On the other hand, vision-language models (VLMs) carry rich semantic knowledge about how fine-grained categories differ from each other, which could complement visual category discovery. Yet existing approaches use this knowledge only through predicted labels or textual descriptions, and a VLM's words are least reliable exactly on the fine-grained novel classes. In this paper, we show that the internal semantic structure of a VLM provides useful supervision beyond its category predictions. We introduce BOLT, a framework that prompts a frozen VLM once per image for a coarse-to-fine taxonomy, reads its hidden state at every level, and uses this multi-level representation as a second view of the pseudo-label hierarchy of the vision encoder: the half of the contrastive embedding that GCD methods leave to an unsupervised loss is trained to the learner's hierarchical pseudo-labels from the image, from the VLM and from their average, and the VLM votes in the clustering that produces those pseudo-labels. BOLT never uses a VLM output as a label and leaves the baseline's objective and evaluation unchanged, so the discovery model keeps its visual discrimination while absorbing the VLM's knowledge. On three fine-grained benchmarks, added to five GCD methods (AL-GCD, SelEx, Hyp-SelEx, SimGCD and CMS) and four backbones (CLIP and DINOv1/v2/v3), BOLT improves AL-GCD by 5.7 points on average and by up to 9.3, and improves the DINOv1-based baselines by 11.5 on average and by up to 31 points, showing the benefit of the complementary semantic knowledge of VLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.