acceptodds
Under review as a conference paper at ICLR 2027

Visual Structure-Guided Textual Semantic Space Adaptation for Domain-Shifted Generalized Category Discovery

Abstract

Domain-Shifted Generalized Category Discovery (DS-GCD) requires categorizing an unlabeled target domain containing both source-seen and target-private categories under domain shift. Existing VLM-based approaches commonly cluster target visual features and represent discovered categories with anonymous visual prototypes, leaving the textual space unused as the target assignment coordinates. We investigate an alternative division of roles: target visual geometry provides unlabeled structural evidence, whereas text embeddings define the coordinates of the target assignment space. We propose Visual Structure Guided Textual Semantic Space Adaptation (ViSTA), which uses target visual structure to search a predefined noun vocabulary for textual anchors and then adapts CLIP within the resulting textual semantic space. In Stage 1, ViSTA filters candidate nouns and greedily selects anchors according to the compactness and separation of the target partitions they induce, under a given target category budget. In Stage 2, the selected words remain fixed while source supervision and target regularization adapt both CLIP encoders. Across four DS-GCD benchmarks and three seen/unseen splits, ViSTA attains a 75.76% average H-Score over 12 settings. These results support textual anchors as category coordinates while retaining target visual geometry as selection evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.