Recovering Evidence Sufficiency Beyond Retrieval Similarity
Abstract
A retrieval system returns documents for every query, whether or not the answer is there. We would like to tell, from the embeddings alone, when the retrieved evidence is actually sufficient, before spending a generator on it. Prior work suggests this cannot be read off the geometry: fixed signals barely beat cosine similarity, so the geometry looks blind. We show the blindness was in the test, not the geometry. Those tests used an additive probe on the two vectors, which cannot represent any interaction between them and therefore cannot even represent cosine; it was asked to find a signal in a language that could not express it. Restore the interaction and the signal reappears, on all nine encoders we test and on a second, structurally different entailment corpus. Two facts explain the earlier null. The signal is not held in any few directions but spread redundantly across many, which is exactly why a fixed cosine readout misses it while a probe that pools directions recovers it. And the probe is nearly free: one dot product over embeddings the retriever already holds, orders of magnitude below a cross-encoder. That is cheap enough to gate every query before generation, with a finite-sample guarantee on how often sufficient evidence is wrongly rejected.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.