acceptodds
Under review as a conference paper at ICLR 2027

Residual Retrieval Memory for Forgetting Decisions in Hybrid RAG

Abstract

Deciding whether a hybrid retrieval-augmented generation system has forgotten a fact requires testing the distributed state that can reconstruct it, not only the disappearance of one direct answer. On Natural Questions, eight repairs occupy a narrow 0.8–3.1% direct range while spanning 8.2–16.7% paraphrase and 10.4–21.3% multi-hop leakage, producing different stopping decisions from nearly indistinguishable direct scores. We formalize the missing decision object as residual retrieval memory and introduce the Residual Retrieval Access Benchmark (RRA-Bench), which profiles six reconstruction channels, separates query-only attacks from target-bearing reintroduction, and pairs fact compromise with retained utility and repair cost. A crossed study of 150 method–stage–seed cells finds that direct leakage has Spearman [0.15, 0.31] with indirect risk and only 0.61 [0.58, 0.65] AUROC for reintroduction-augmented AnyLeak@32; at a 1% direct threshold, 13 of 18 passing cells remain false-safe. PathSweep, a trace-conditioned allocator over dense, sparse, metadata, reranker, and cache surfaces, provides a constructive test under the same operations and budgets as fixed and validation-selected schedules; it lowers Hard Delete's reintroduction-augmented AnyLeak@32 from 38.6% [36.2,41.1] to 18.7% [16.4,21.2] while retaining 0.793 global MRR. Disjoint prompts, independently authored Qwen-2.5 attacks, 72-hour routine ingestion, and a component-replaced HotpotQA stack preserve the residual-risk separation hidden by direct scores. Residual retrieval memory thus preserves the distinctions needed to stop and rank hybrid-RAG repairs after direct answers disappear.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.