acceptodds
Under review as a conference paper at ICLR 2027

Uncovering the alignment and repelling tradeoff in CLIP: a path towards Gaussian geometry

Abstract

The contrastive learning framework introduced by CLIP learns image and text representations in a shared embedding space, aligning corresponding samples while repelling representations across the dataset. Building on the alignment–uniformity perspective of Wang and Isola, we identify the term responsible for uniformity in the CLIP objective as a repelling term. We show that its optimum is incompatible with perfect cross-modal alignment: improving the repelling term beyond the uniform distribution on the sphere necessarily requires a positive paired gap. To address this, we propose a prototype objective that shifts repulsion from individual embeddings to representations of image-text pairs. We prove that, under a mild condition on the inverse temperature, its unique population minimizer is the uniform distribution on the sphere with perfect cross-modal alignment. More generally, when perfect alignment is not possible, the optimal distribution is uniform on a sphere whose radius is determined by the HGR maximal correlation. We further show that fixed-dimensional projections of the optimal representations converge to an isotropic Gaussian as the embedding dimension increases. Experiments on MS-COCO and CC3M show that fine-tuning CLIP with the prototype objective substantially reduces the modality gap and improves spherical uniformity and coordinate Gaussianity, while maintaining comparable probe performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.