acceptodds
Under review as a conference paper at ICLR 2027

Same Explanations, Opposite Effects: How Data Construction Shapes WebPRM Judgment

Abstract

Web process reward models (WebPRMs) use explanations to help agents choose actions. When training labels influence how explanations are generated, models may learn to recognize labels rather than judge actions. We introduce PreAct, which reconstructs both candidates' explanations in WebPRM Collection using one generator and only pre-execution information, without access to preference labels. To evaluate how models use explanations, we introduce WebContrast, a benchmark of 300 pages across 119 websites that tests correct responses to goal changes and stability across explanation versions. On WebContrast, the same test explanations hurt action judgment on average after original-data training but help after PreAct training. Relative to each model's accuracy without explanations, their effect shifts from -3.58 to +5.58 percentage points for Qwen3-8B and from -24.25 to +5.33 for Llama-3.1-8B-Instruct. PreAct also improves mean action-selection accuracy on WebLINX and WebShop, and on Collection with pre-execution or no explanations. Yet evaluation with Collection's original explanations favors original-data training. Our findings show that data construction can change whether explanations help or hinder action judgment, while evaluation using the original explanations can obscure this difference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.