What Do Negatives Buy? A Controlled Anatomy of Contrastive vs. Negative-Free Text Embeddings
Abstract
Contrastive negatives are the default for unsupervised sentence embeddings; what do they buy? We swap SimCSE's InfoNCE for a negative-free Gaussianity penalty with view prediction, holding encoder, data, views, pooling and projector fixed. This prices two tuned packages, not one term; we also intervene inside InfoNCE. On RoBERTa, negatives buy a trade-off rather than a uniform advantage: at development-selected checkpoints, they raise semantic similarity (STS) and cost topic clustering. The cost appears at every corpus size in our main study; without selection it attenuates at our largest corpus. At our smallest corpus its sign holds in every architecture clearing our competence bar, decisively in all but the smallest. Under RoBERTa's main recipe the STS benefit grows with training scale, which one pass ties to corpus size. Fresh pre-registered seeds replicate this growth in sign, with no learning-rate dependence demonstrated; a pre-registered rerun at a lower negative-free learning rate did not demonstrate the growth. BERT, at optima from a one-seed search, already shows the benefit at our smallest corpus. On RoBERTa at matched count, dropping same-topic rather than random negatives raises clustering, even with labels from no neural model. Our pre-registered headline, that InfoNCE degrades at small batch, was not supported.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.