Residual Perturbations as Sentence-Pair Generators for Contrastive Sentence Embedding Training
Abstract
Sentence embeddings, which map sentences to dense vectors that capture their meaning, underpin many NLP tasks such as semantic search, clustering, and retrieval. Training sentence embeddings typically relies on sentence pairs, which are costly to generate. To solve this resource bottleneck, researchers have developed alternative methods to generate sentence pairs efficiently. Unsupervised alternatives bypass this by forming positive pairs from single sentences. SimCSE is a well-established unsupervised method that uses dropout noise to create two views of the same sentence as a positive pair. In this article, we introduce and explore residual perturbations as an alternative sentence pair generation method. Residual perturbations generate sentence pairs by adding perturbations onto the original embedding between the embedding layer and subsequent attention layers, treating the resulting perturbed embedding as the positive counterpart to the original. We investigate different embedding-level perturbations: random noise, rolling, shuffling, and adding the sentence's centroid. The results highlight that embedding-level residual perturbations performs on par with dropout-based methods, such as SimCSE. On the MTEB benchmark, residual rolling and residual shuffling perturbations on average match or slightly exceed SimCSE after unsupervised contrastive pretraining, with residual shuffling winning the largest share of individual tasks. Adding the centroid recovers much of this benefit through a simple constant shift, while random noise and, more severely, discarding the original embedding underperform even an untouched baseline. While our method is comparable to SimCSE overall, viewing performance per task category reveals clear differences in performance with narrow losses on retrieval and clustering tasks and strong improvements on STS and pair classification tasks. The strong performance of the residual shuffle perturbation on STS tasks suggests that perturbations can be chosen according to the downstream task at hand, offering an additional layer of control over sentence embedding generation, as well as a working alternative to dropout-based methods of sentence pair generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.