Revealing and Employing Oblique Semantic Geometry in Vision Foundation Models
Abstract
Vision foundation models learn rich representation spaces that encode a wide range of visual semantics. Yet, understanding the organization of this semantic structure often relies on predefined concepts, supervision, or training additional models. Thus, it remains unclear which semantic factors emerge directly from the representation distribution itself, and how these factors relate to one another. In this work, we first examine, in a controlled manner, the relations among semantic attributes in known datasets. We observe that similar geometry emerges across both vision-only and vision-language encoders, suggesting that this structure is largely data-driven. Since semantic directions are generally non-orthogonal, we develop a geometric framework for analyzing their organization. To recover the underlying factors without labels or semantic supervision, we seek a complete linear decomposition that separates distinct sources. Independent Component Analysis (ICA) provides a natural construction for this purpose by separating statistically independent, non-Gaussian factors through a complete and invertible transformation. Its dual representations form a biorthogonal pair, enabling precise feature disentanglement while exposing the geometry between semantic factors. We show across several measures and datasets that ICA recovers semantic factors more effectively than PCA and SAEs, with less identifiable sources being harder to recover. We further find that geometrically aligned factors form coherent semantic groups, revealing structure from fine-grained attributes to broader concepts. Finally, we employ this structure through controlled interventions, demonstrating its utility for image retrieval and mitigating spurious correlations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.