acceptodds
Under review as a conference paper at ICLR 2027

AHEAD: Attention Heads for Adversarial Detection for GUI Agents

Abstract

Graphical-user-interface (GUI) agents rely on grounding models to translate screenshots and element descriptions into interaction coordinates, making grounding a point of exposure to visual adversarial attacks. Existing detectors require adversarial fitting examples or repeated inference, motivating detection using signals available during grounding inference. We observe that visual attacks alter how attention is distributed over screenshot regions during coordinate prediction, producing measurable changes in attention entropy. Motivated by this observation, we propose AHEAD (Attention HEads for Adversarial Detection), which jointly measures head-wise entropy deviations from a clean reference. AHEAD extracts its features during a single grounding inference and uses Mahalanobis distance with a threshold calibrated on clean inputs, requiring no adversarial fitting examples or additional model inferences. Across four grounding models, two interface datasets, and three attack settings, AHEAD achieves a mean AUROC of and a mean true positive rate of at a matched false positive rate (TPR@FPR), with the lowest overhead compared to the baselines. AHEAD also transfers from mobile to web interfaces and vice versa using only clean inputs from the new domain, achieving a mean TPR@FPR in both directions. Overall, AHEAD provides a lightweight safeguard between visual perception and action execution, enabling intervention before unintended interactions and supporting safer deployment of GUI agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.