TraceEdit: Self-Evolving Multi-Channel Environment Attacks on GUI Agents
Abstract
GUI agents built on multimodal large language models can be hijacked by adversaries who mutate the agent's environment, and reported persistence of that hijack across planning, memory, and tool-use stages is a recurring empirical finding. We argue that this finding is partly an artifact of measurement: when the adversarial trace is still on the screen during evaluation, observed persistence conflates continued exposure with an internalized shift in the agent's behavior, and the latter remains underexplored as a training objective. Our claim is that persistence should be defined and trained counterfactually, by removing the adversarial trace before re-evaluating the agent. We operationalize this claim with TraceEdit, a self-evolving attacker that maintains a reversible edit trace over five environment injection surfaces (grouped into channels: text, visual saliency, layout, interaction flow, and event timing) and is trained under a constrained reinforcement-learning objective whose persistence term is computed by a deterministic counterfactual re-evaluation in which every reversible edit has been undone. Multi-channel coordination is the setting in which rollback-conditioned persistence has the most room to differ from exposure, and a dual variable enforces a fixed detectability constraint on every emitted edit. Empirical results across four GUI-agent targets and six security benchmarks demonstrate that TraceEdit outperforms a broad set of baselines on attack effectiveness, residual persistence, and stealth, and that its rollback-conditioned persistence is empirically separable from on-screen-time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.