Protocol Rendering Shapes Source Boundaries in LLM Agents
Abstract
LLM agents combine task instructions and external content in structured inputs. We study how Protocol Rendering, defined by wrapper syntax, field order, key style, and nesting, affects attack susceptibility when the receiver model, task, and attacker payload are held fixed. To analyze and investigate this effect, we introduce protocol-conditioned source separability (PCSS), which measures how distinctly a model represents identical neutral text placed in trusted versus external fields. We perform controlled experiments across five models and three attack settings and conclude three important findings : (1) Rendering alone changes attack outcomes. Switching from XML to nested JSON produces a 31% relative increase in attack success on Qwen3-14B memory-attack tasks, yet no wrapper family is consistently safest. (2) Weaker source separation is associated with higher memory-attack risk across renderings in three models spanning two model families. (3) Shifting external-message representations toward a trusted-source direction learned from neutral prompts without attack labels increases attack success by 9.2% on held-out Qwen3-14B memory-attack tasks, demonstrating a causal influence of source representations on attack susceptibility. These findings establish Protocol Rendering as a security-relevant design choice and support source separation as a contributor to rendering-dependent attack risk. https://anonymous.4open.science/r/ProtocolRenderingSafety-2848Code is available online.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.