SAGE-RAG: Defending Retrieval-Augmented Generation with Trust Calibration and Graph Selection
Abstract
Retrieval-augmented generation (RAG) augments language model inputs with retrieved evidence to improve performance on knowledge-intensive tasks. When attackers can inject content into the retrieval corpus, however, this reliance on external evidence creates opportunities to manipulate generated answers. We study coordinated evidence poisoning, in which individually plausible and relevant passages reinforce misleading claims that conflict with benign evidence. Such attacks motivate assessing relations among claims, since local plausibility and apparent agreement alone do not establish trustworthiness. We propose SAGE-RAG, which represents retrieved claims and their relations in an evidence graph and combines trust calibration with joint evidence selection. Attack-Aware Trust Propagation (AATP) uses structural coordination patterns as signals to calibrate claim trust. Maximum Consistent Reasoning Subgraph selection (MCRS) then jointly selects claims under a token budget by balancing utility, calibrated trust, and relational consistency. It downweights claims with low trust and penalizes the joint selection of contradictory claims. On an 8,000-sample benchmark spanning four QA datasets and four attack types, SAGE-RAG reduces attack absorption from with BM25 to and achieves the highest exact-match (EM) score among ten evaluated systems, reaching compared with for RobustRAG. Additional evaluations show gains across four generators and on the RGB benchmark. Ablations show that trust calibration helps retain benign evidence, while joint selection further reduces the inclusion of adversarial claims.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.