acceptodds
Under review as a conference paper at ICLR 2027

What Do Attack-Success Rates Measure? A Matched-Sham Control for GUI-Agent Injection Benchmarks

Abstract

Prompt-injection evaluations of graphical user-interface (GUI) agents often count whether an agent acts on an injected instruction. Such decision compliance is distinct from completing an attacker’s objective, and a malicious-only score cannot show how often an agent would act on a benign instruction delivered in the same way. We use a three-arm comparison—malicious, matched sham, and clean—to measure this distinction under a fixed execution scaffold. On 57 live VPI-Bench browser-use cases, sham instructions were followed in 52/114 first post-exposure actions and 84/114 episodes at any step, compared with 112/114 and 114/114 for malicious instructions; the clean arm was 0/114. In a preregistered page-realism-by-interactivity analysis, sham engagement was 25/80 on toy pages under one-shot probing, 24/740 on toy pages under interactive probing, and 0/119 and 0/120 in the two real-page cells. In email and messenger tasks, the malicious-versus-sham contrast was inconclusive after symmetric rescoring. These results describe decision compliance under our tested scaffold; their rate differences and ratios do not decompose the task-completion rates reported by the original benchmark.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.