Reasonable-RLVR: Step-Verified Rewards for Reasoning over Knowledge Graphs
Abstract
Reinforcement learning with verifiable rewards (RLVR) trains reasoning models with programs that check their outputs. Such programs check answers in mathematics and code, and they check steps in formal theorem proving. Knowledge-intensive domains such as medicine have no such language, so a correct answer reached by unsound reasoning receives full credit. Knowledge graphs make entailment decidable, but no reward checks cited steps against their closure. We introduce Reasonable-RLVR, which verifies each cited proof step against the OWL 2 RL closure of a graph and checks whether the steps establish the answer. The verifier tests the supported chain against a relation contract, has configurable coverage, and needs no reference proof or learned judge. We define two versions of the reward, a two-band form that satisfies five ordering properties on anchored traces and a simpler weighted sum. Reasonable-RLVR obtains a higher correct-and-valid rate than outcome and reference-path rewards on four knowledge-graph benchmarks. Without a supervised warm-up on WebQSP, it reaches 0.597 ± 0.120 correct-and-valid proofs over three seeds, while the outcome reward reaches 0.000 at 94% answer accuracy. Verifier feedback at inference also improves frontier-model proofs. These results extend verifiable rewards to reasoning over knowledge graphs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.