acceptodds
Under review as a conference paper at ICLR 2027

RIPPLE: Corpus-scale Retrieval Plan Optimisation via Iterative Failure-driven Supervision

Abstract

Retrieval-augmented generation (RAG) increasingly relies on iterative search when evidence cannot be recovered in a single retrieval pass. Existing iterative and RL-based search policies reformulate queries across search rounds, but fall short in certain aspects: the subqueries and documents considered per round bound how much of the corpus is ever brought into view, and although the planner adjusts its queries round to round, it does so mainly by reading its own accumulated search history within its context, with no trainable signal for how well a round resolved the information need, leaving little basis to learn which decisions closed the evidence gap. We introduce RIPPLE, which casts a wide initial retrieval net and compresses it with maximal marginal relevance (MMR) into a compact, diversity-optimized snapshot that anchors the planner's corpus view from the first round. Conditioned on this snapshot and prior search history, the planner issues subqueries; an LLM judge evaluates retrieved evidence and summarizes it as feedback for the next round. We train the planner with Group Relative Policy Optimization (GRPO), judging the answer generated from evidence gathered at each round to produce a reward at the level of individual rounds rather than only at the episode's end. On MSQA, TechQA, and ANTIQUE, RIPPLE improves LLM-based scores by up to 7% and CAR by up to 15% over strong baselines, including Search-R1, with ablations confirming both the diverse snapshot and search-round-level supervision are independently necessary. These gains generalize across planner backbones, an independent evaluator, multi-hop benchmarks (HotpotQA, 2WikiMultiHopQA), and zero-shot transfer, while also reducing search rounds needed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.