acceptodds
Under review as a conference paper at ICLR 2027

RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation

Abstract

We introduce RAG-Stress, a controlled diagnostic protocol for examining the limits of evidence reliance in retrieval-augmented generation. Misleading retrieved evidence can induce a model to replace an answer it previously gave correctly, a failure that aggregate accuracy alone does not distinguish from preexisting errors. The protocol holds the question and reference answer fixed, edits one assertion to support a designated incorrect answer, and crosses two source priority policies with three positions of the answer span within the evidence text. We measure misleading rate (MR) on each model's subset of questions answered correctly without retrieval, alongside clean accuracy on the full evaluation set. We evaluate fifteen systems spanning API models, open models, and search agents trained with reinforcement learning on TriviaQA-RC, HotpotQA, and SearchQA, with additional English and Chinese MedQA evaluations. Instructions that prioritize documents consistently produce higher MR than those permitting reliance on prior knowledge. Averaged over models and positions, the gap ranges from 10.9 to 13.5 percentage points across the three QA datasets. Averaged equally across models, MR follows End Beginning Middle under both policies on all three QA datasets; this ordering is not universal across individual models. A separate paired audit of 500 questions and two checkpoints supports increased harmful override under Strict relative to Soft, while the corresponding instruction contrasts in beneficial correction remain inconclusive. These findings distinguish evidence adherence from factual reliability and motivate evaluating whether retrieved evidence preserves, replaces, or corrects a model's answers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.