acceptodds
Under review as a conference paper at ICLR 2027

Valid Actions, Wrong Targets: Visual Attacks and Evidence-Based Defense for Web Agents

Abstract

Web agents translate natural-language requests and webpage observations into browser actions, while visual web agents additionally use screenshots to capture information that page text and structure may omit. This visual channel can also be manipulated: an agent may be redirected by screenshot content even when non-visual evidence is sufficient to determine the correct next action. We investigate this risk across three web-agent interfaces using paired tasks and controlled visual interventions, complemented by web observations. Image masking and image–text conflicts show that visual reliance changes with the task, even on the same page. We evaluate local noise, global noise, single-patch, and multi-patch attacks on open-weight models through generated native actions and target-element binding. Motivated by these findings, we propose Evidence-Grounded Action Repair (EGAR), a training-free defense that separates task understanding, local evidence extraction, and evidence-based correction. EGAR bypasses the screenshot when non-visual evidence uniquely determines a supported action. Otherwise, it reads task-relevant attributes from local candidate regions and checks them against the task requirements before selectively revising the target element. Offline evaluations show improved recovery of adversarially redirected actions and correction of some clean-task errors, while preserving most correct clean decisions. Our findings show that controlling visual exposure and explicitly checking task-relevant evidence provide complementary ways to improve the robustness of visual web agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.