SYNERGATE-RAG: Learning Graph-Grounded Process Rewards for Multi-Hop Question Answering
Abstract
Multi-hop retrieval-augmented generation requires a reader to connect facts across evidence sources, yet final-answer supervision provides little information about the validity of individual reasoning steps. A trajectory with a correct final answer can therefore receive full credit despite unsupported intermediate claims or disconnected evidence use. Knowledge graphs provide explicit factual constraints, but incomplete coverage limits their reliability as process supervision. We introduce SynerGate-RAG, which learns continuous graph utility from source matching, citation relevance, and validated contradiction witnesses. Unlike graph process reward models (PRMs) based on algorithmic correctness or implicit outcome signals, our framework learns explicit, source-grounded utility and controls its contribution to graph-text process rewards. The labels preserve partial support and separate unavailable evidence from contradiction. An evidence-conditioned PRM learns this utility, while a stepwise gate combines graph and text rewards under a source-availability mask. The resulting rewards guide reader training alongside answer correctness. Experiments on multi-hop question answering benchmarks show improved answer quality and support the complementary roles of reward learning and adaptive allocation. SynerGate-RAG thus turns incomplete structured knowledge into selective process supervision without reward-model calls at inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.