acceptodds
Under review as a conference paper at ICLR 2027

Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG

Abstract

Retrieval-augmented generation (RAG) grounds large language models (LLMs) in external sources. However, retrieved passages often name the right entities while lacking the actual facts an answer requires. Generators rarely notice: even with an explicit instruction to abstain, 12 generators still answer 40.0–99.3% of insufficient-evidence questions. A common remedy trains abstention into the generator through fine-tuning or reinforcement learning. This ties the decision to one model's weights, potentially rewarding correct answers recalled from its parametric knowledge as if the evidence had supported them. The decision arrives only as the generator's own output, so every insufficient-evidence question still costs a full generator call. We ask instead whether sufficiency can be judged from the question and evidence alone, before any answer exists. We first demonstrate several pitfalls in constructing tests for insufficient evidence. Directly removing the relevant evidence or pairing evidence with unrelated questions can introduce cues that reveal the label without testing sufficiency, such as lexical overlap and evidence position. We therefore build a paired benchmark from substitution, deletion, and question-swap constructions that vary answer support while controlling selected surface features, such as word use. On this benchmark, sufficiency can be judged without generating an answer, but no single signal works on every dataset. We introduce RINSE (Relevance Is Not Sufficient Evidence), which combines three signals, one for each way evidence falls short: whether every part of the question is covered, whether any passage actually offers an answer, and whether a small language model, reading the passages together, treats them as sufficient. Across six datasets, RINSE ranks sufficient above insufficient evidence with a score of 0.837 (chance 0.5), ahead of the best of 10 prior methods (0.746) and of a frontier model asked the same question through an API (0.784). Its weakest dataset scores higher than any other method's weakest (0.684 vs. 0.676). It runs locally before generation, taking only 36.5 ms per question on a single GPU. Our code is available at https://anonymous.4open.science/r/RINSE-reproduction/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.