Resource Attribution in RAG: Selection, Verification, and Utilization
Abstract
We develop a resource attribution framework for recovery after limited fact verification in retrieval-augmented generation. Grounded in statistical decision theory and information-channel comparisons, the framework separates ideal information advantages from rendering loss and a fixed generator's utilization gap. It distinguishes privileged selection of existing evidence from acquiring verification messages, and derives matched controls for attributing observed recovery. A 200-group study on HotpotQA and MuSiQue instantiates these controls across content availability, claim status, and answer selection. Same-text verification under competition raises reference-alias EM by 28.5 and 34.0 points for two Qwen generators, with a smaller, dataset-dependent interaction for Mistral. Retrieved-document controls reveal a different attribution: matches rise from 104/200 to 135/200 when an already-visible oracle-selected quotation is re-exposed as ordinary evidence, then to 137/200 when it is marked verified. The incremental EM effect is 1.0 point [], with secondary benefits in F1 and replacement matching. A frozen answer-blind selector reaches 98/200, failing to retain the oracle gain. These results demonstrate how the framework separates protocol-specific recovery effects and limits their extrapolation, while keeping reference recovery distinct from appropriate abstention. To facilitate further research in this area, we release our code anonymously at https://anonymous.4open.science/r/fsws62dwas212fdf345.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.