Structured Dots in the Sentence Embedding Space
Abstract
Embedding spaces allow us to manipulate texts as mathematical objects, most frequently as points or vectors paired with distance metrics as proxies of linguistic properties, mainly similarity or relatedness. Viewed through this perspective, the embedding space appears anisotropic, and thus not able to encode linguistic differences accurately enough, despite empirical proof that the embeddings are useful for a variety of tasks. We propose that these observed shortcomings do not reflect properties of the embeddings themselves, but of the shallow measures such as cosine, which consider each dimension separately. By comparing the relative positions in the sentence embedding space of three sentence representation variations, and their performance on a variety of tasks, we show that sentence embeddings are complex objects with internal structure, which encodes linguistic structure in a systematic manner.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.