ReGain: Information-Gain Guided Rollback for Retrieval-Augmented Reasoning
Abstract
Large language models can learn to interleave reasoning with search engine calls through reinforcement learning, enabling multi-step retrieval-augmented generation. However, existing approaches either rely on sparse outcome-only rewards with no per-turn signal, or even when providing process-level supervision, retain misleading retrieval content in the trajectory without correction. We propose ReGain, a closed-loop search quality control framework that integrates two complementary components: (1) a per-turn quality signal based on the change in the model's answer confidence before and after each retrieval, which serves as both a rollback trigger and a step-level process reward; and (2) a state rollback mechanism that detects and undoes low-quality searches in real time during RL rollouts, restoring generation state and injecting explicit query reformulation feedback. Experiments on seven QA benchmarks show that ReGain achieves the highest average F1 among existing outcome-reward and process-reward RL methods, with particularly strong gains on multi-hop reasoning tasks
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.