Do LLMs Consider Document Credibility When Relying on Retrieved Evidence?
Abstract
Evaluations of retrieval-augmented generation (RAG) typically examine whether large language models extract retrieved evidence, resolve conflicts, and detect when evidence is insufficient. These settings assume that the retrieved document is trustworthy, while prior work on credibility relies on external cues such as source metadata. We ask instead: when a retrieved document provides evidence whose veracity the model cannot determine, but contains credibility flaws elsewhere, does the model reduce its reliance on that evidence? We call this capability credibility-aware evidence reliance and introduce CREDO (Credibility-aware Reliance Evaluation on Documents) to measure it. CREDO pairs real retrieval-dependent questions with documents that retain the target evidence but carry five types of intrinsic credibility flaws. The truth of the target evidence is unknowable to the model by construction, so document credibility is the only rational basis for reliance. Measured with four response-level behavioral metrics, most open-weight and proprietary models adopt the flawed document's claim in 85% to 99% of responses and identify credibility flaws in fewer than 10%, and more capable models are not more prudent. Furthermore, frequent search calls co-occur with low flaw identification, pointing to habitual tool use rather than credibility assessment. Prompting yields limited gains, whereas fine-tuning on synthetic evidence-reflection trajectories improves credibility-aware evidence reliance and performance on general retrieval-based QA tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.