acceptodds
Under review as a conference paper at ICLR 2027

AsymStego: Exploiting Information Asymmetry to Bypass Content-Level Auditing in LLM Agents

Abstract

Indirect prompt injection attacks exploit untrusted content processed by LLM agents. Content-only auditors can screen this content before it reaches the agent, but the auditors and the executing agent may operate under different runtime information states, such as different timestamps, causing the same content to be interpreted differently. We investigate how this information asymmetry can be exploited to evade content-only auditing. We present AsymStego, an environment-bound steganographic injection attack that embeds instruction fragments in natural-language carrier text and enables the agent to recover them using runtime information that differs from that available to the auditor. We evaluate AsymStego on 62 attack cases derived from InjecAgent attacker intentions in a controlled, task-driven stress test of content-level auditing. On the OpenClaw–DeepSeek-V4-Flash stack, a GPT-5.5 auditor reduces baseline attack success rates from 83.9%–92.5% without auditing to 0.5%–19.4%. Under the same permissive audit, when the attacker has knowledge of the agent's relevant environment, AsymStego achieves a 52.7% attack success rate, compared with 90.3% without auditing. These results suggest that content-only auditing can face challenges in such settings, motivating runtime-aware defenses that incorporate the agent's execution context.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.