acceptodds
Under review as a conference paper at ICLR 2027

DualView: Preventing Indirect Prompt Injection in Computer Agents

Abstract

Personal AI agents that run on the user's local machine, such as OpenClaw, automate tasks by accessing the network, file system, and shell. This access exposes them to indirect prompt injection (IPI) attacks. Prior Dual LLM defenses prevent attacker instructions from steering tool-call decisions by replacing untrusted data with symbols that the agent can reference but not read. However, they track untrusted data only inside the agent's context, so an attacker's prompt can return as trusted data after the agent saves and rereads it, causing stored IPI. We present DualView, which extends untrusted data tracking from the agent's context to the local files while preserving original data for humans and non-agent programs. In AgentView, the agent sees untrusted data as symbols even after writing and rereading it, which blocks stored IPI. HumanView exposes original data to humans and non-agent programs. DualView routes tool calls and synchronizes data across the two views. DualView deploys as a plugin using only tool hooks, without changing the agent's tool-call logic or tool implementations. DualView deterministically prevents instructions in untrusted data from directly steering the agent's tool calls, and this guarantee is not limited to the evaluated attack templates. In our evaluation on an IPI benchmark and PinchBench, DualView blocked every tested instruction-injection IPI attack, including stored IPI, while utility on PinchBench dropped by 1.8 and 6.4 percentage points from the unprotected baseline on Claude Haiku 4.5 and Claude Sonnet 4.6, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.