acceptodds
Under review as a conference paper at ICLR 2027

Understanding Generative Contrastive Learning via the Conditional Generation Kernel

Abstract

Recent advances in unsupervised text embedding leverage the stochasticity of autoregressive generation by LLMs to generate semantically related positive pairs for contrastive learning. The resulting paradigm, which we term generative contrastive learning, has the unique property that the data-generating distribution is explicitly parametrized by the LLM, thus opening the door to a formal analysis of the contrastive objective. Building on this, we introduce the Conditional Generation Kernel (CGK), a positive semi-definite kernel that measures textual similarity through expected co-generation probabilities across a set of anchor documents. We show that the standard InfoNCE loss optimizes dot-product similarities toward the logarithm of the CGK, which exposes a fundamental representational bottleneck: because the logarithm of a kernel matrix is generally indefinite, the optimal target cannot be exactly realized as an inner product in any Euclidean space. Motivated by this bottleneck, we derive a contrastive objective whose global optimum is exactly the CGK and find that it aligns with the known spectral contrastive loss. We empirically confirm that the two objectives actually approach their theoretically derived optimal states, resulting in clearly differing geometries of the resulting embedding spaces. Finally, we evaluate the resulting models on downstream tasks, demonstrating that training with the spectral loss yields favorable properties for clustering and classification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.