Duet-PRM: Factorizing Plan Reliability and Step Correctness in Process Reward Models
Abstract
Process reward models (PRMs) provide fine-grained supervision for multi-step reasoning, but most existing approaches collapse different sources of reasoning failure into a single step-level reward. We identify this issue as plan-step conflation: an incorrect trajectory may arise either from an unreliable plan or from incorrect execution of an otherwise reasonable plan. To address this, we propose Duet-PRM, a generative PRM that separately predicts plan-level reliability and step-level correctness. We construct the required factorized supervision through a fully automated pipeline based on multiple independent executions, without human annotation or auxiliary LLMs. Across four mathematical reasoning benchmarks, the two reward signals provide complementary information, and their combination consistently improves chain discrimination and downstream Best-of- selection over either component alone. Duet-PRM also achieves the strongest average candidate-selection performance among the evaluated PRMs. Further analysis shows that higher predicted plan scores consistently correspond to higher reasoning success rates, and that plan-level signals can correct selection errors when step-level evidence alone favors an incorrect answer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.