ForkPRM: Learning Embodied Process Verifiers from Executable Counterfactuals
Abstract
Process supervision improves language-model reasoning, but step-level correctness labels are hard to obtain for embodied agents because an action's quality depends on its downstream environmental consequences. We introduce ForkPRM, which turns that dependence into a supervision signal: during training, candidate actions are executed from the same observed navigation history and each resumes the same frozen continuation policy, so the measured utilities of the resulting branches give baseline-relative process labels without human annotation. These labels train a candidate-set verifier that scores candidates at application time and selectively overrides or defers to the policy action without executing additional environment forks. On R2R vision-and-language navigation, shared-prefix supervision raises ordinary pairwise ranking over unpaired outcome regression in the same five-seed supervision panel and label budget (+2.4 pp, positive in all five seeds; +7.6 pp over a cross-state permutation control), but these ranking gains do not extend to utility-weighted accuracy, rescue recall, or harm rate. A one-shot evaluation on scenes reserved from the verifier's splits—drawn, like the policy's own training, from the R2R training-environment pool—yields a small branch-selection gain whose confidence interval includes zero against a much larger budget-matched oracle; ranking accuracy is therefore not yet policy utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.