acceptodds
Under review as a conference paper at ICLR 2027

Explorer-R1: Efficient Iterative Retrieval via Stepwise Evidence-State Rewarding

Abstract

Iterative retrieval enables large language models to gather evidence step by step, which a single search cannot provide. Existing iterative retrieval methods have used reinforcement learning (RL) to improve multi-hop RAG retrieval for question reasoning and answering. However, most approaches are end-to-end optimizations that use the final answer as the reward signal, which misses intermediate reasoning rewards and makes efficient training difficult; moreover, the model cannot serve as an independent retrieval module for other models. To solve these issues, we propose Explorer-R1 (ExR1), an RL framework for training a plug-and-play multi-hop RAG retrieval model. It reformulates multi-hop retrieval as a sequential decision process. Specifically, we decompose the original end-to-end process into a sequence of state transitions. By tracking evidence state transitions, we collect data to optimize subsequent model decisions. During the RL stage, we design a dense state-level reward that enforces format legality, evidence pruning, and search-result alignment to maintain a compact evidence state. ExR1 decouples retrieval from answering, enabling it to serve as an independent, robust retrieval module that adapts seamlessly and flexibly to downstream collaborative tasks with other large models. Extensive experiments on HotpotQA, MuSiQue, and 2WikiMultihopQA show that ExR1 outperforms existing iterative RAG baselines in recall, converges faster during training, and can effectively collaborate with different answer models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.