When Semantically Equivalent Contexts Yield Different Answers: Measuring Contextual Uncertainty to Improve Retrieval-Augmented Question Answering
Abstract
A fundamental challenge of retrieval-augmented generation (RAG) is that retrieving relevant evidence does not guarantee that an LLM uses it consistently or correctly. Current RAG systems lack a model-agnostic mechanism to identify such failures and use the diagnosis to improve QA performance. We introduce SECURE (SEmantic Context Uncertainty for REtrieval-based QA), a training-free method requiring only text input/output access to the generator. SECURE rephrases one retrieved chunk at a time, then measures the semantic consistency of the resulting answers across these intended meaning-preserving context variants. The resulting score estimates answer risk, while each single-chunk comparison exposes which edit changed the response, without requiring reference answers or model internals. Across four benchmarks, SECURE achieves state-of-the-art performance in uncertainty evaluation, using the same correctness criterion for all methods. We then close the diagnostic-to-improvement loop: SECURE-RAG uses this same signal to trigger adaptive semantic chunking and guide iterative context refinement. On the selected high-uncertainty questions, SECURE-RAG improves reported QA accuracy in 15 of 16 backbone–benchmark combinations. Together, these results show that evidence-rephrasing sensitivity can both rank answer risk and guide targeted RAG refinement in the evaluated settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.