acceptodds
Under review as a conference paper at ICLR 2027

CrossView-PRM: Falsifiable Step-Level Credit from Correlated Noisy Views

Abstract

Reinforcement learning for long-horizon language agents commonly requires verifiable outcomes or externally supervised process reward models (PRMs), assumptions that fail in open-ended planning with multiple valid solutions and correlated model judges. We introduce CrossView-PRM, a target-verifier-free reward-construction framework, with three main contributions. Rather than treating agreement as correctness, the framework uses unlabeled target statistics only to detect instability and relies on source diagnostic domains to test whether learned reliability rules transport. Matched-budget sibling actions turn each decision into a controlled counterfactual comparison, yielding a non-telescoping ordinal reward without target-domain labels. Semantic, model, and temporal views are residualized and redundancy-weighted, while a source-calibrated gate abstains under view collapse, disagreement, or domain shift; a conditional analysis bounds accepted reward-sign error. Controlled causal planning, TravelPlanner, and TripScore jointly test reward validity, transport, policy utility, robustness, and cost. CrossView-PRM reaches \(0.45\) Kendall's \(\tau\), \(76%\) pairwise accuracy, and \(30%\) TravelPlanner Final Pass, improving over Majority-view PRM by \(0.12\) \(\tau\) and \(8.5\) final-pass points. Shared-bias stress tests and resource accounting further characterize the reward's reliability and cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.