acceptodds
Under review as a conference paper at ICLR 2027

SYNERGATE-RAG: Learning Graph-Grounded Process Rewards for Multi-Hop Question Answering

Abstract

Multi-hop retrieval-augmented generation requires a reader to connect facts across evidence sources, yet final-answer supervision provides little information about the validity of individual reasoning steps. A trajectory with a correct final answer can therefore receive full credit despite unsupported intermediate claims or disconnected evidence use. Knowledge graphs provide explicit factual constraints, but incomplete coverage limits their reliability as process supervision. We introduce SynerGate-RAG, which learns continuous graph utility from source matching, citation relevance, and validated contradiction witnesses. Unlike graph process reward models (PRMs) based on algorithmic correctness or implicit outcome signals, our framework learns explicit, source-grounded utility and controls its contribution to graph-text process rewards. The labels preserve partial support and separate unavailable evidence from contradiction. An evidence-conditioned PRM learns this utility, while a stepwise gate combines graph and text rewards under a source-availability mask. The resulting rewards guide reader training alongside answer correctness. Experiments on multi-hop question answering benchmarks show improved answer quality and support the complementary roles of reward learning and adaptive allocation. SynerGate-RAG thus turns incomplete structured knowledge into selective process supervision without reward-model calls at inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.