acceptodds
Under review as a conference paper at ICLR 2027

ClickShield: A Firewall for Multimodal Computer-use Agents Against Dual Prompt Injection

Abstract

Multimodal computer-use agents execute goals by repeatedly interpreting a screenshot and a natural-language instruction. Because today’s agents implicitly trust both the user prompt and on-screen content, they are vulnerable to dual prompt injection: malicious instructions and deceptive UI elements (e.g., phishing pop-ups) that steer the agent into unsafe actions. We introduce ClickShield, a training-free, model-agnostic runtime firewall that wraps an existing agent without retraining or modifying its policy. ClickShield enforces a zero-trust boundary via two parallel modules: (i) an instruction risk filter using an LLM-as-a-judge, and (ii) a decoupled UI parser that detects interactable regions, extracts OCR text, and produces lightweight element captions to form a structured, DOM-like screen representation. A judge LLM then compares the verified goal against this element list to flag intent-conflicting elements; ClickShield deterministically blackout-masks their bounding boxes before the agent acts. Across RiOSWorld-Bench (492 risky tasks) and OS-Harm-Bench, ClickShield sharply suppresses attack success (e.g., adversarial goal completion (59.6% 5.4%) while preserving benign utility on OSWorld (only marginal success degradation).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.