acceptodds
Under review as a conference paper at ICLR 2027

CORE: Internalizing Context Faithfulness via Intervention-Derived Process Rewards in Reinforcement Learning

Abstract

In context-grounded tasks, Large Language Models may underutilize provided contextual information and instead rely more heavily on their parametric knowledge. Although reinforcement learning can optimize models toward task-specific objectives, outcome-based rewards provide only sparse supervision and offer limited direct guidance on how the context should be used during generation. Process-based rewards provide finer-grained supervision over intermediate generation steps, but often rely on costly intermediate annotations. We propose CORE, a reinforcement learning framework for improving contextual faithfulness during generation. Given a trajectory generated with access to the context, CORE re-evaluates the likelihood of the same generated tokens after intervening on context availability and uses the resulting token-level likelihood changes to measure contextual influence. This signal is incorporated as a dense, annotation-free process reward during policy optimization. To mitigate reward exploitation through excessive context copying, CORE further introduces an accuracy-gating mechanism that applies the contextual influence reward only to correct trajectories, coupling contextual reliance with task correctness. Across multiple model architectures, reinforcement learning algorithms, and both in-distribution and out-of-distribution benchmarks, CORE consistently improves task accuracy over the corresponding baselines. Complementary LLM-as-a-Judge evaluations further indicate improved contextual faithfulness on the evaluated settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.