Preserving Fine-Grained Structure under Coarse Supervision via EMA-Guided Similarity-Weighted Contrastive Learning
Abstract
Fine-tuning a pretrained model with coarse labels can destroy existing fine-grained structure in latent space. Preserving this latent structure is important when detailed annotations are unavailable. Standard supervised contrastive learning (SupCon) amplifies this problem by attracting all samples sharing a coarse label uniformly, despite potentially substantial within-class heterogeneity. We introduce , a similarity-weighted extension of SupCon that uses a stop-gradient exponential moving average (EMA) teacher to estimate within-class pairwise similarity and selectively redistribute positive-pair attraction. Similar positives therefore receive greater learning pressure, while dissimilar positives are attracted less strongly, without requiring fine-grained labels, an auxiliary self-supervised objective, or an external memory bank. We evaluate TaxoCon on iNaturalist 2021 and FGVC-Aircraft using only coarse labels for representation fine-tuning and model selection, with finer-grained labels withheld for post-hoc evaluation. Across ResNet-50, ViT-S/16, and Swin-T pretrained backbones, TaxoCon consistently improves preservation of fine-grained structure across neighborhood retrieval, linear-probing, and hierarchy-sensitive metrics on evaluated baselines. These results show that TaxoCon can mitigate the loss of fine-grained structure during coarse-supervised fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.