Relevance Says Yes, but Sufficiency Says No: Efficient Adaptive Retrieval via Latent Evidence Sufficiency
Abstract
Adaptive retrieval-augmented generation (RAG) must decide whether additional retrieval is needed, which evidence is worth acquiring, and when retrieval should stop. Existing approaches typically rely on retrieval relevance, uncertainty, or explicit model judgments, yet relevance does not indicate whether the current evidence is sufficient, and highly relevant passages may still be redundant. We further find that explicit sufficiency judgments can be systematically misaligned with the sufficiency information encoded in hidden representations. Building on these observations, we introduce **SuREST**, an adaptive RAG framework that uses a lightweight probe to recover latent evidence sufficiency and employs it as a unified control signal throughout retrieval. Across four multi-hop QA benchmarks, SuREST achieves the best answer quality among the evaluated baselines. With Qwen2.5-7B-Instruct, it improves macro F1 by 2.36 points over IRCoT, the strongest baseline by macro F1, while reducing retained passages by 56.7% and end-to-end inference time by 53.6%. Evidence analysis suggests that these gains come from reducing redundant evidence while largely preserving coverage. These results demonstrate that latent evidence sufficiency provides an effective basis for jointly improving retrieval effectiveness, evidence quality, and efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.