acceptodds
Under review as a conference paper at ICLR 2027

NC-CLIP: Geometric Feature Decoupling for Robust Vision-Language Alignment

Abstract

Contrastive pre-training dictates modern vision-language alignment, yet its conventional "align-everything" objective is fundamentally misaligned with the nature of multimodal data. By forcing the holistic alignment of images and texts, existing frameworks conflate shared semantics with modality-specific noise. This leads to the memorization of spurious correlations, the destruction of complementary information, and fragile generalization under distribution shifts. We propose that robust representation requires explicit decoupling: we must align the shared semantic core while retaining modality-specific traits. To this end, we present , a framework that mathematically enforces this separation via the geometry of Neural Collapse. leverages bidirectional feature masking alongside a fixed Equiangular Tight Frame (ETF) that acts as an absolute geometric anchor. This design strictly guides the alignment of pure semantic features while isolating modal noise. Extensive benchmarking demonstrates that drastically improves robustness to distribution shifts and spurious correlations, all while maintaining competitive zero-shot performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.