acceptodds
Under review as a conference paper at ICLR 2027

SciPRI: Process-Reward Inversion in Scientific Agents

Abstract

Process reward models (PRMs) score the intermediate steps of language-model agents and increasingly drive step selection, tree search, and reinforcement learning. Almost all public PRMs are trained on mathematical reasoning, where a step’s value is its local arithmetic/logical correctness. Scientific agents (data analysis, experiment design, ML engineering) instead produce steps whose value is the information they contribute toward a correct scientific conclusion, often carried by terse, information-dense actions (code, observations, exploratory tests) rather than fluent prose. We introduce SciPRI, a measurement-and-mitigation framework that quantifies per-step information gain 〖VIG〗_t with a calibrated probe over frozen hidden states (the change in the calibrated log-probability of task success) and uses it to audit reward models. Our central finding is process-reward inversion: the most widely used PRMs, math-trained discriminative reward models, become negatively aligned with information gain on scientific steps. The same model flips sign between domains: Qwen2.5-Math-PRM has length-controlled partial correlation ρ=+0.209 and mixed-effects slope β_I=+0.0094 (p< 10^(-29)) between reward and information gain on a math control, but ρ=−0.170, β_I=−0.0101 (p=3 × 10^(-42)) on scientific tasks. The inversion is gold-validated on 102 ScienceAgentBench tasks (math PRMs correlate −0.132 to +0.045 with gold scientific step quality, where they should be strongly positive), is robust across six generator models, three alternative observer models, and many estimator choices, and harms Best-of-N selection below random. A panel of reward models reveals a spectrum of failure: discriminative math PRMs invert, value-head PRMs collapse, and a strong general LLM-critic mis-selects. We introduce InfoPRM, a lightweight reward head trained to track information gain and predict the outcome while remaining length-decorrelated; it matches the best Best-of-N result on BLADE (0.711) and wins outright on ScienceAgentBench (0.810 vs. 0.749), agrees with gold scientific step labels (+0.294), and, in reward-guided generation, improves a weak agent’s task success by 16.1–18.9 percentage points. SciPRI argues that process supervision for scientific agents must be measured against information gain, not against resemblance to mathematical reasoning in actual scientific-agent practice.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.