Answering by Format: How Evidence Formatting Shapes LLM Abstention in RAG
Abstract
We study evidence-grounded answering and abstention in retrieval-augmented generation (RAG), where the decision to answer should depend on whether the evidence is sufficient. However, we identify a spurious evidence-rendering gate: changing how the same evidence is presented, without altering its content or order, can switch a large language model (LLM) between answering and abstaining. In controlled experiments with nine LLMs, correlating evidence format with whether training examples call for an answer or abstention induces this gate in eight of them. On insufficient evidence, this training increases the answer-rate gap between the two formats by an average of 0.70. Reversing this association reverses the learned format bias in most models. To assess the performance consequences of this association, we vary its strength and find that accuracy on the answer-or-abstain task declines as the association strengthens. To examine whether such correlations also occur in practice, we audit public training corpora and find associations between formatting cues and answer-versus-abstention labels. Controlled retraining further shows that artificially introducing such correlations can induce the gate. Beyond these controlled experiments, we also observe format-dependent answering in open (ChatQA-1.5-8B; TrustAlign-Qwen2.5-3B/7B) and closed (claude-opus-5) models. In the controlled experiments, removing the association between evidence format and answer-versus-abstention labels improves task accuracy by up to 0.37 relative to format-correlated training. Reliable abstention in RAG therefore requires controlling evidence presentation during both training and evaluation. We release our code at https://anonymous.4open.science/r/paper-artifact-staging-B0B7/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.