acceptodds
Under review as a conference paper at ICLR 2027

When Step Boundaries Change the Reward: Auditing Process Reward Model Specifications

Abstract

Process reward models (PRMs) score reasoning steps, but completed solutions need not have a unique step partition. We formalize partition well-definedness and introduce PRM-PART to audit candidate selection while fixing the normalized ordered payload and moving only contiguous boundaries. Across three open PRMs, boundaries alter decisions beyond mechanical aggregation. In two separately frozen 1,000-problem cohorts, source-to-canonical parsing and frozen multi-view scoring change 13.2–14.7% and 25.4–43.9% of tolerance-robust winners for two PRMs; the multi-view result also transfers across two rollout generators. A selected Qwen audit retains 25/420 pools after text and AI semantic screens. Under a subsequent model-card direct-join control, boundary changes alter 13/25 winners, 11 robustly; the equal-step/equal-token subset has 9/19 changes, 7 robust. Fresh score-only search produces an actual wrong selection in one of 38 eligible eight-candidate Qwen pools (128 screened) in a separately frozen Level-5 stratum, with CPU confirmation. This stratum was motivated by the original sample's answer composition. The original 33/128 eligible pools show no harm; Level 5 also contains a repair, leaving Qwen accuracy unchanged. Equal-budget random search finds the same harmful case. Together, these results establish partition-dependent selection and an attainable correctness failure under fixed content. Reproducible scoring therefore requires a partition contract specifying normalization, atomization, ownership, serialization, readout, and aggregation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.