CrossView-PRM: Falsifiable Step-Level Credit from Correlated Noisy Views
Abstract
Reinforcement learning for long-horizon language agents commonly requires verifiable outcomes or externally supervised process reward models (PRMs), assumptions that fail in open-ended planning with multiple valid solutions and correlated model judges. We introduce CrossView-PRM, a target-verifier-free reward-construction framework, with three main contributions. Rather than treating agreement as correctness, the framework uses unlabeled target statistics only to detect instability and relies on source diagnostic domains to test whether learned reliability rules transport. Matched-budget sibling actions turn each decision into a controlled counterfactual comparison, yielding a non-telescoping ordinal reward without target-domain labels. Semantic, model, and temporal views are residualized and redundancy-weighted, while a source-calibrated gate abstains under view collapse, disagreement, or domain shift; a conditional analysis bounds accepted reward-sign error. Controlled causal planning, TravelPlanner, and TripScore jointly test reward validity, transport, policy utility, robustness, and cost. CrossView-PRM reaches \(0.45\) Kendall's \(\tau\), \(76%\) pairwise accuracy, and \(30%\) TravelPlanner Final Pass, improving over Majority-view PRM by \(0.12\) \(\tau\) and \(8.5\) final-pass points. Shared-bias stress tests and resource accounting further characterize the reward's reliability and cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.