Same Explanations, Opposite Effects: How Data Construction Shapes WebPRM Judgment
Abstract
Web process reward models (WebPRMs) use explanations to help agents choose actions. When training labels influence how explanations are generated, models may learn to recognize labels rather than judge actions. We introduce PreAct, which reconstructs both candidates' explanations in WebPRM Collection using one generator and only pre-execution information, without access to preference labels. To evaluate how models use explanations, we introduce WebContrast, a benchmark of 300 pages across 119 websites that tests correct responses to goal changes and stability across explanation versions. On WebContrast, the same test explanations hurt action judgment on average after original-data training but help after PreAct training. Relative to each model's accuracy without explanations, their effect shifts from -3.58 to +5.58 percentage points for Qwen3-8B and from -24.25 to +5.33 for Llama-3.1-8B-Instruct. PreAct also improves mean action-selection accuracy on WebLINX and WebShop, and on Collection with pre-execution or no explanations. Yet evaluation with Collection's original explanations favors original-data training. Our findings show that data construction can change whether explanations help or hinder action judgment, while evaluation using the original explanations can obscure this difference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.