acceptodds
Under review as a conference paper at ICLR 2027

Answerless Bridges: Relevant to Whom? One Model Needs Them, Another Does Without

Abstract

Retrievers for language models are trained and evaluated with relevance labels: a passage is relevant when it contains the answer, resembles the question, or is annotated as supporting evidence. Each label reads only the question and the passage, so it gives every model the same verdict and never asks: relevant to whom? Multi-hop questions make this concrete, because a model can depend on an answerless bridge, a passage without the answer that links the question to the passage containing it. For each model, we delete either the bridge or a control passage from the same context and call the bridge necessary when deleting it alone breaks a correct answer. In each of eleven open-weight models on HotpotQA and MuSiQue, bridge deletion breaks more answers than control deletion, yet for most questions evaluated with several models, the same bridge is necessary for some models and dispensable for others. Lexical, dense and cross-encoder relevance scores give these models one value and predict need no better than chance. ContextCite attribution, fitted to each model's responses to context ablations, predicts need in all ten models tested, and on a shared question it separates the models that need the bridge from those that do not, partly through each model's base rate. The same test recovers the operands of DROP arithmetic questions, whose answers appear in no sentence, and on synthetic prompts shown identically to base and instruction-tuned models it shows that post-training strengthens dependence on the passages holding the answer. The answer to "relevant to whom?" is a particular model, whose own responses measure its need.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.