AppRAG: When Relevant Evidence Does Not Apply in Retrieval-Augmented Generation
Abstract
Retrieval-augmented generation (RAG) uses retrieved text passages as evidence to answer knowledge-intensive questions. However, relevant and factually correct evidence may not satisfy the conditions of the reasoning step it is used to support. When the context can include only a limited number of passages, applicable evidence may rank too low to be selected. Meanwhile, individually useful passages may not jointly support all necessary reasoning steps. To address these evidence-selection challenges, we introduce AppRAG, a three-stage framework for selecting applicable and complementary evidence. First, requirement-conditioned applicability modeling assesses whether evidence supports question-derived requirements under matching conditions, distinguishing valid partial support from condition mismatch while keeping unresolved evidence separate. Second, applicability-guided embedding calibration uses these assessments to update the query representation and rerank available candidates, verifying evidence replacements before updating the context. Third, applicability-aware evidence selection constructs complementary evidence sets under a per-context document budget and compares the answer drafts generated from each set. The answer generator remains frozen throughout. Our analysis characterizes when missing condition-match information imposes unavoidable evidence-selection loss. On evaluation sets from HotpotQA, MuSiQue, and 2WikiMultiHopQA, AppRAG improves Qwen3-8B's overall token-level F1 by 2.82-13.43 pp across all eight tested RAG integrations, including a 10.61-point gain over Vanilla RAG. Our code is available at https://anonymous.4open.science/r/AppRAG-C2E8
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.