Recursive Dense Retrieval
Abstract
*Recursive retrieval* comes in a few flavors: *single-hop* retrieval iteratively expands context around a single concept, whereas *multi-hop* retrieval focuses on *depth*, where it iteratively continues an evidence chain, and/or *width*, searching for complementary pieces of evidence. Although all scenarios require, at each retrieval step, decisions on (i) what to search for next and (ii) how to organize the evidence gathered along the way, both encoder training and inference strategies largely remain scenario-specific. In this work, we introduce **RDR**, a recursive dense retrieval framework that unifies single- and multi-hop retrieval. At the core of RDR is an instruction-conditioned encoder trained to map a retrieval goal and observed artifacts into a dense next-step query. Repeated encoding and index lookups build a retrieval tree, and we produce readouts (i.e., evidence rankings) by aggregating evidence scores across the tree's nodes. Because the instruction is part of the embedded state, the same encoder can both refine single-hop queries and follow multi-hop evidence chains. We train an RDR-8B encoder on **Recursive Retrieval Worlds** (RRW), a 4.75 million-sample dataset compiled from existing retrieval data, contextual transformations, constructed worlds, and recovery supervision. RDR performs strongly on a wide range of single- and multi-hop evaluations. In multi-hop retrieval, with a four-level retrieval tree with a branching factor of three, RDR reaches 80.14 and 25.28 Recall@5 on MuSiQue and BrowseComp+, respectively, outperforming state-of-the-art encoders and multi-hop retrievers on all four multi-hop benchmarks. In single-hop retrieval, RDR's tree-based readout improves the encoder's five-benchmark mean by 2.5 points, most on conversational retrieval (70.79 nDCG@10) and on multi-aspect reasoning queries, where it nearly closes the gap to the strongest reasoning retriever. On average, RDR outperforms the strongest evaluated baselines in both task families, substantially on multi-hop retrieval (by 5.8 Recall@5) and by 3.6 nDCG@10 on single-hop retrieval. We will publicly release RDR-8B, RRW data, and the evaluation code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.