acceptodds
Under review as a conference paper at ICLR 2027

ReGain: Information-Gain Guided Rollback for Retrieval-Augmented Reasoning

Abstract

Large language models can learn to interleave reasoning with search engine calls through reinforcement learning, enabling multi-step retrieval-augmented generation. However, existing approaches either rely on sparse outcome-only rewards with no per-turn signal, or even when providing process-level supervision, retain misleading retrieval content in the trajectory without correction. We propose ReGain, a closed-loop search quality control framework that integrates two complementary components: (1) a per-turn quality signal based on the change in the model's answer confidence before and after each retrieval, which serves as both a rollback trigger and a step-level process reward; and (2) a state rollback mechanism that detects and undoes low-quality searches in real time during RL rollouts, restoring generation state and injecting explicit query reformulation feedback. Experiments on seven QA benchmarks show that ReGain achieves the highest average F1 among existing outcome-reward and process-reward RL methods, with particularly strong gains on multi-hop reasoning tasks

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.