DGPO: Directional Gain Policy Optimization for Credit Assignment in Reasoning
Abstract
Outcome-level rewards supervise entire reasoning trajectories, even when their success depends on a few intermediate segments. We present Directional Gain Policy Optimization (DGPO), which converts binary outcome rewards into information-weighted segment-level credit without process-level supervision. DGPO constructs uncertainty-triggered reasoning trees to expose alternative continuations under shared prefixes. Its core credit signal, Directional Information Gain (DIG), combines normalized information gain with a signed measure of success association, capturing both how strongly a segment predicts the outcome and whether it is associated with success or failure. Global DIG captures tree-wide associations, while local DIG contrasts sibling branches; their normalized scores augment trajectory-level advantages for policy optimization. Across four language-model backbones and six mathematical reasoning benchmarks, DGPO achieves higher average pass@1 accuracy than GRPO, TreePO, and TreeRL on every backbone. With the same tree-construction procedure, DIG also outperforms direct segment advantages and TreeRL advantages in a four-benchmark comparison on Qwen3-4B-Base. A separate comparison on this backbone shows higher final accuracy at comparable total training time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.