acceptodds
Under review as a conference paper at ICLR 2027

Multi Surface Attack: An Agent Is Only as Safe as Its Weakest Surface

Abstract

Large language model (LLM) agents expose numerous attack surfaces, e.g., user messages, system prompts, tool outputs, environment states etc. Yet existing studies typically inject malicious prompts through only one specific surface or a limited combination of surfaces. These studies generally bind each malicious prompt to a fixed attack surface. In this paper, we systematically study how the attack surface affects success by delivering the same attack through different surfaces of an agent. We further study whether a chat LLM is vulnerable to such agentic-style multi-surface attacks when provided with a synthetic agentic context payload. Finally, we study what happens when an agent or LLM is attacked through multiple surfaces simultaneously. We evaluate nine common agent attack surfaces across five large language models and four existing benchmarks. Our extensive experiments show that routing the same malicious prompts through a more vulnerable surface can increase the attack success rate (ASR) by up to 100% relative to the baseline surface. We also show that chat LLMs are similarly vulnerable to agentic multi-surface attacks when we use fake agent-like context payloads. We find that attacking multiple surfaces simultaneously can significantly increase the ASR. However, attacking more surfaces does not necessarily yield a higher ASR and can even make an agent more resistant to attacks, as surfaces can interact both positively and negatively. Based on our findings, we propose a novel multi-surface attack (MSA) approach for any agent or chat LLM that can increase the attack success rate by up to 100%. Finally, we release a comprehensive multi-surface attack benchmark constructed by adapting attacks from four existing benchmarks using our methodology.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.