Automedon: Taking the Reins of Web Agents through Full-Adaptive Manipulation
Abstract
Large language model-based web agents increasingly execute user tasks through autonomous browser interactions, but their reliance on webpage content exposes them to prompt injection. Existing attacks are zero-adaptive or partial-adaptive. They fix payloads in advance or adapt only to limited observed states, and thus fail under the execution dynamics of long-horizon tasks. We present Automedon, the first full-adaptive manipulation framework for black-box web agents. Automedon builds on two insights. (i) Step-to-step webpage transitions form an observable footprint of the agent's decision-making, and (ii) attacker-controlled browser extensions provide a persistent attack surface for modifying what the agent observes. To steer an agent through a long-horizon malicious task, Automedon coordinates a Planner, a Generator, and an Evaluator for explore-exploit planning, state-conditioned payload generation, and transition-based verification, respectively. Across four datasets, Automedon achieves an average ASR of 0.99, outperforming seven baselines while preserving benign task accuracy within 0.04 of the optimal value. Experimental results also demonstrate that it remains effective across different backend LLMs, observation functions, agent frameworks, and countermeasures. Our code is available at https://anonymous.4open.science/r/Automedon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.