acceptodds
Under review as a conference paper at ICLR 2027

What Lies Within: Diagnostics of Compositional and Sequential Behavior in Long-Context Text Embeddings

Abstract

Embedding evaluation has long relied on indirect external tasks such as retrieval, owing to the opaque, vector-valued nature. We depart from this dependence and present a paradigm that directly analyzes the compositional behavior of the output vector under cumulative synthesis. **RSS** measures the coherence between a text and its embedding representations, while **SRS** measures the order sensitivity across positions in text. **RSS** correlated with downstream retrieval performance (**LEMB**, Pearson , and on NarrativeQA task). **SRS** likewise correlated with **LEMB**, with the negative **SRS** score correlating at . **RSS** required relatively little measurement time compared with **SRS** and **LEMB** ( that of **LEMB**); it can therefore be used to analyze performance by sequence length quickly. Every embedder exhibited a distinguishable **SRS** pattern, and these patterns tracked each model's fine-tuning lineage and base model. Measuring and visualizing order sensitivity-as **SRS** does for single-vector embedders-constitutes, to the best of our knowledge, the first attempt to interpret and structure how the position and order information of a document is preserved at the final representation level. FINESSE, the framework behind these results, will be released as a Python package.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.