acceptodds
Under review as a conference paper at ICLR 2027

Better Contrast Decisions, Worse Ranking? How Validation Selects Readouts for Text Embeddings

Abstract

A scoring function can rank a passage above its negated counterpart yet place it below other documents in a collection. We study how to select functions that improve this contrast while retaining the benchmark answer's ranking performance. Using frozen text embeddings, we hold each candidate bank fixed within a selection experiment and vary the documents and criteria used for validation. On NevIR, inspection after selection finds candidates that improve contrast accuracy with no decrease in measured test ranking. Yet contrast-only validation repeatedly selects functions with substantial ranking losses. Including retrieved competitors and prioritizing ranking quality reduces these losses while retaining contrast gains. Competitor identity matters even at a fixed list length, and more labeled contrasts do not replace missing background comparisons. The same selected set can also improve average fixed-list reranking while lowering average full-corpus ranking. On NFCorpus, the trained candidates offer only small gains in distinguishing relevance grades, leaving little improvement for a selector to recover. These results distinguish gains available among the trained candidates from gains recovered by validation, and show why a selected scoring function must be evaluated in its intended retrieval setting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.