acceptodds
Under review as a conference paper at ICLR 2027

The Price of Imperfect Process Rewards: Bias Floors and Finite-Budget Reversals

Abstract

A learned process reward model (PRM) provides feedback at each step, whereas its reward estimates may be inaccurate, which can lead to suboptimal policies. We compare policy learning from a fixed PRM with learning from unbiased outcome feedback under the same downstream rollout budget. On task families where the PRM preserves most preferences, we establish a reversal through matching bounds: process feedback has lower minimax regret at small positive budgets, while outcome feedback has lower minimax regret at large budgets. With an unknown transition model, the minimax regret of the process learner decreases toward an exact bias floor. Our lower bound accounts for the transition estimation error, establishing the reversal even when actions have an arbitrarily small effect on transition probabilities. In the known-transition family, for trajectories of steps, the crossover budget scales as for the largest process regret over a range of per-stage reward gap. For finite tabular MDPs with compact convex occupancy sets, we give exact bounds on the loss in cumulative true reward under policy comparisons with a specified tolerance in PRM value. These bounds account for potential shaping and reward scale, and policy constraints can make these class-specific certificates arbitrarily smaller than the corresponding global STARC guarantees. For a weighted combination of process and outcome rewards, the bias term in the regret bound is proportional to the PRM weight. Our tabular experiments vary PRM accuracy at a fixed true reward, examining how horizon length and transition estimation affect policy regret.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.