acceptodds
Under review as a conference paper at ICLR 2027

Trust or Abstain? A Self-Aware RAG Approach under Knowledge Conflicts

Abstract

Retrieval-augmented generation (RAG) improves large language models (LLM) by incorporating external context, but it also introduces knowledge conflicts when retrieved contextual knowledge (CK) and parametric knowledge (PK) disagree, or both fail. Existing approaches mainly coordinate which source to use, without explicitly asking whether each answer path is correct. We argue that faithful RAG requires LLM self-awareness, namely the ability to recognize the limits of its own knowledge and reasoning. To ground this problem, we construct a model-specific, ground-truth-aligned knowledge-conflict benchmark by evaluating four LLM backbones on PK-only and CK-conditioned answer paths over approximately 69K query-context instances per backbone, drawn from five conflict-QA datasets. We then introduce SABER, a lightweight Self-Aware Belief Estimator for RAG. SABER extracts and combines a self-prior and source-conditioned reasoning representations from frozen LLM hidden states, estimates PK-side and CK-side correctness beliefs with two lightweight predictors, and maps them into a 4-cell decision space covering trust PK, trust CK, trust both, or abstain. Across four backbones and five datasets, SABER improves end-to-end accuracy and conflict-specific faithfulness over ten baselines, with the largest gains on conflict-heavy datasets. When configured to abstain, SABER's risk-coverage curve Pareto-dominates every prompt-based abstainer, providing a tunable balance between coverage and risk on answered.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.