acceptodds
Under review as a conference paper at ICLR 2027

CTRL-RAG: Contrastive Likelihood Reward Based Reinforcement Learning for Context-Faithful RAG Models

Abstract

With the growing use of Retrieval-Augmented Generation (RAG), improving contextual faithfulness and retrieval-grounded reasoning in large language models (LLMs) has become increasingly important. Existing RAG-oriented reinforcement learning (RL) methods primarily rely on external reward signals, such as correctness or LLM-based judges, which often provide indirect or noisy supervision for document faithfulness. Meanwhile, internal self-reward signals based on model confidence lack explicit grounding in retrieved evidence and may reinforce spurious behaviors during training. To address these limitations, we propose CTRL-RAG, an internal–external hybrid reward framework centered on a Contrastive Likelihood Reward (CLR). CLR measures the log-likelihood difference between responses generated with and without supporting evidence, directly estimating the contribution of retrieved documents to generation. This encourages the model to identify and utilize relevant evidence while suppressing reliance on noisy context. Combined with a correctness reward through a gating mechanism, CTRL-RAG encourages the model to use relevant evidence while treating correctness as a prerequisite. Experiments on single-hop, multi-hop, biomedical, and faithfulness benchmarks show consistent improvements in retrieval-grounded reasoning and contextual faithfulness across both dense and Mixture-of-Experts architectures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.