Calibrated Category Discovery: Knitting Calibration from the Known into the Novel
Abstract
Generalized Category Discovery (GCD) aims to recognize old classes while discovering novel ones in unlabeled data, yet the practical value of such discoveries hinges on whether the model's confidence can be trusted. Existing GCD methods optimize clustering accuracy alone, leaving their predictions on novel classes markedly miscalibrated. We argue that this deficiency cannot be resolved by conventional post-hoc calibration. First, miscalibration in GCD is inherently cluster-level, whereas conventional methods apply a single sample-level mapping to all clusters. Second, novel classes provide no labels for calibrating confidence. To address this, we formulate Calibrated Category Discovery (CalCD) and introduce a cluster-conditional expected calibration error (cluECE) that assesses calibration within each predicted cluster, preventing the deviations of different clusters from cancelling as they do in the standard expected calibration error. We further propose Known-to-Novel calIbration Transfer (KNIT), grounded in the premise that the trustworthiness of a cluster may follow a class-agnostic pattern. KNIT characterizes cluster reliability through the model's self-assessment and an independent semantic check by a vision-language model, estimates this pattern on labeled old-class clusters, and transfers it to novel-class clusters, correcting confidence without altering predictions. Extensive experiments comprising 220 comparisons across 4 GCD methods, 11 datasets spanning generic, fine-grained and multi-domain images, and 5 calibration baselines demonstrate that KNIT significantly reduces cluECE for both novel and old classes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.