acceptodds
Under review as a conference paper at ICLR 2027

Rewarding the Graph Behind the Chain: Topology-Aware Process Supervision for RL and Reasoning Distillation

Abstract

Long-form reasoning is produced as a linear token sequence, even when a conclusion depends on multiple non-adjacent premises. Outcome-only rewards do not expose this support structure, while scalar process scores do not explicitly identify which earlier premises support a later conclusion. We introduce TopoPRM, a hierarchical process reward model that recovers typed support graphs from ordinary reasoning text. The same dependency representation drives two forms of supervision: asymmetric credit estimation (ACE) assigns structural credit under final-answer correctness constraints, and topology-guided distillation (TGD) directs a fixed teacher to revise the current student's reasoning. Fresh student rollouts close the loop between structural diagnosis and online supervision. The deployed policy generates ordinary text without graph decoding or a verifier. Across four mathematics benchmarks, the complete pipeline improves pass@1 by up to 13.3 percentage points over outcome-only training while reducing mean response length by 3.8%-24.0%. Against the distillation-matched w/o ACE control, it gains 2.6 points in average pass@1 and uses 14.8% fewer tokens. These results connect correctness-constrained structural credit with more compact reasoning, using a shared representation for reward and targeted revision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.