Grounded-R1: Reinforcement Learning for Preemptive Hallucination Intervention and Grounded Generation
Abstract
Hallucination in large language models (LLMs) often begins locally: an ungrounded claim can enter the generation trajectory, steer subsequent reasoning, and compound into a global failure. Retrieval-augmented generation (RAG) can reduce such failures by supplying external evidence, but to prevent local error propagation, evidence must arrive before risky claims propagate. However, existing methods do not directly learn retrieval trigger timing from local hallucination risk. Retrieval is driven either by heuristic rules or fixed procedures in training-free methods or by delayed, coarse indirect rewards in reinforcement learning (RL)-based approaches, causing evidence to arrive too late to prevent local error propagation while unnecessary calls waste the retrieval budget. In this paper, we propose Grounded-R1, a dynamic RAG framework for preemptive hallucination intervention. Grounded-R1 introduces the Real-time Hallucination Risk Detection (RHRD) module to estimate sentence-level hallucination risk from generation-time signals and uses two-phase RL to align retrieval trigger timing with this risk. Across eight knowledge-intensive benchmarks, Grounded-R1 improves final-answer exact match (EM) by 18.0% over the strongest external baseline with comparable retrieval calls. On five open-ended benchmarks, it raises sentence-level non-hallucination rate (NHR) by 12.7 points. Delivering the same evidence one sentence earlier reduces next-sentence hallucination rate by 11.1 points; under equal retrieval budgets, Grounded-R1 targets positions with an observed 8.3-point downstream hallucination reduction. Grounded-R1 thus improves grounded generation not by retrieving more, but by retrieving where one sentence of delay matters. Our code and models are available at https://sites.google.com/view/grounded-r1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.