Reward on Path: Learning Intermediate Supervision Signals for Knowledge Graph Question Answering
Abstract
Knowledge Graph Question Answering (KGQA) aims to answer user questions by reasoning over Knowledge Graphs (KGs). Recent methods use supervision derived from answer labels or refined by Large Language Models (LLMs) to train models that retrieve KG evidence for LLM-based answer reasoning. However, answer-derived supervision treats every answer-reaching path as correct and thus yields noisy training signals, whereas LLM-refined supervision mitigates this noise at substantial cost. To address these limitations, we propose Reward on Path (RoP), a framework to learn a lightweight, question-conditioned path reward from answer labels with an asymmetric objective. Paths reaching the same answer are supervised jointly as a bag, allowing the reward model to learn their relative contributions, while each path is penalized individually for retrieving non-answer entities. The learned reward then trains an LLM-based relation path generator in two stages: reward distillation transfers reward-induced preferences over candidate paths into the generator, and on-policy optimization with GRPO further refines the policy on self-generated paths. Generated paths are grounded in the KG to retrieve evidence for answer reasoning. Experiments on multiple KGQA datasets show that RoP improves F1 over answer-derived methods by at least 2.6% while outperforming LLM-refined methods without costly supervision construction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.