Beyond Local Positives: LLM-Verified Diverse Positive Mining for Generalized Category Discovery
Abstract
Generalized Category Discovery (GCD) aims to recognize known categories and discover novel ones from unlabeled data given limited labeled examples of known categories. A central difficulty is constructing effective training signals for unlabeled instances, especially for novel categories. Recent LLM-assisted methods primarily improve the semantic reliability of these signals, but under-emphasize positive diversity: the coverage of same-class pairs that are less similar under the current representation. An oracle study with ground-truth filtering confirms this gap: even when all positives are same-class pairs, restricting them to highly similar candidates still limits discovery. However, expanding the candidate pool also introduces false-positive relations when ground-truth labels are unavailable. To address this trade-off, we propose DiPo (Diverse Positive mining), a budgeted two-path framework for neighborhood contrastive learning. For a fraction of anchors, DiPo expands the candidate pool and uses LLM same-or-different judgments to retain semantically valid, less-similar positives; for the remaining anchors, it uses the nearest cluster-consistent mutual neighbor as a reliable local positive. Extensive experiments show that DiPo attains the best average performance among strong LLM-assisted baselines and is particularly effective on novel-class discovery. Code is available at https://anonymous.4open.science/r/DiPo-5244DiPo.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.