acceptodds
Under review as a conference paper at ICLR 2027

WebPACED: Evaluating Web Agents under State-Contingent Risks

Abstract

Web agents are increasingly deployed to perform consequential actions in everyday applications. A significant but underexplored risk is that faithfully completing a benign request can cause harm, even in a non-adversarial environment. We term this a state-contingent risk, characterized by (i) isolated benignness, where neither the instruction nor the environment is risky on its own, and (ii) state dependence, where harm depends on the current application state. Violation rates alone cannot distinguish restraint from task failure or indiscriminate refusal, nor reveal whether an agent's execution responds to application state. We introduce WebPACED, a benchmark of 1,000 native web task instances from 180 task families, spanning five evidence-demand levels that vary how the evidence determining permission is organized. Each protocol fixes the instruction, action, and content while independently varying permission (whether the action is safe to execute) and completion (whether the requested outcome already holds). Execution when a permitted action's outcome is unmet measures capability, while on tasks the agent can complete, withholding once the action becomes unsafe measures safety and stopping once the outcome holds measures state awareness. Across 12 evaluated models, GLM-5.3-Flash and Claude Opus 5 achieve capability scores of 71% and 65%, but conditional safety scores of 32% and 58%, respectively. Every agent withholds more often once the outcome holds than once execution becomes unsafe alongside a divergent sensitivity to safety.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.