Semantic Coverage Imbalance: A Hidden Representational Bias in Visual Learning
Abstract
Modern visual models achieve strong overall accuracy but fail on specific classes. While prior work attributes these failures to class imbalance or dataset bias, recent work has identified a more fundamental cause, Semantic Coverage Imbalance (SCI), as uneven coverage of class–concept relations in training data. We ask a complementary question: how is uneven semantic coverage reflected in the learned representations and behavior of trained visual models? To address this, we introduce a simple, model-agnostic post-hoc framework that measures semantic representation structure and semantic reliance from concept-based evidence, applicable across datasets, architectures, and concept sources. Across spurious-correlation benchmarks (WaterBirds), attribute-based datasets (CelebA), general object recognition (CIFAR-100), context-shift benchmarks (NICO++), and real-world medical imaging (ISIC-DICM-17K, MILK10k), we find that the model-level effects of uneven semantic coverage vary across architectures and domains, are associated with worst-class performance and generalization gaps, and are not explained by class imbalance or dataset-level bias alone. Under shortcut and context shift, failures depend on which semantic concepts a model relies on, not simply on the overall degree of semantic reliance. We further show that post-hoc analysis can identify unstable semantic dependencies that provide useful intervention targets. In medical imaging, however, clinically meaningful concepts remain comparatively stable across domains, suggesting that robust prediction depends on relying on the right semantics rather than uniformly reducing semantic reliance. Overall, our framework offers a principled approach for analyzing and diagnosing SCI as a representation-level bias in visual learning settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.