acceptodds
Under review as a conference paper at ICLR 2027

SIR: Self-improving Red-teaming for Computer Use Agents

Abstract

Computer-use agents (CUAs) are agents powered by vision-language models (VLMs) that perceive a screen and operate an operating system through mouse, keyboard, and terminal interactions to automate everyday digital tasks. Their exposure to untrusted content creates a risk of indirect prompt injection (IPI), where an adversary embeds instructions in content the agent reads to redirect it toward actions that violate the user's intent. Evaluations based on fixed, hand-written injections may underestimate the risk posed by adaptive adversaries. We present **SIR**, a black-box self-improving IPI framework that (i) composes task-specific injections from a small library of reusable red-teaming principles stated in plain language and (ii) uses an iterative feedback loop to analyze unsuccessful attack trajectories and distill new, named principles that are retained in a shared library and reused across tasks. We target operating-system-level compromise and evaluate outcomes through deterministic checks on filesystem, service, and permission state rather than an LLM judge. An attack counts as successful only when both the adversarial objective and the benign user task are completed in the same execution. We evaluate three frontier CUAs, allowing SIR up to 10 attack attempts per case. It achieves joint attack success rates of 24% on Claude Opus 4.8 and 28% on GeminiĀ 3.5 Flash, compared with 4% and 0%, respectively, for the benchmark's fixed, hand-written injection. The red-teaming principles discovered against one victim also improve attacks against other victims, including a different model family, without additional feedback.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.