acceptodds
Under review as a conference paper at ICLR 2027

Protocol Rendering Shapes Source Boundaries in LLM Agents

Abstract

LLM agents combine task instructions and external content in structured inputs. We study how Protocol Rendering, defined by wrapper syntax, field order, key style, and nesting, affects attack susceptibility when the receiver model, task, and attacker payload are held fixed. To analyze and investigate this effect, we introduce protocol-conditioned source separability (PCSS), which measures how distinctly a model represents identical neutral text placed in trusted versus external fields. We perform controlled experiments across five models and three attack settings and conclude three important findings : (1) Rendering alone changes attack outcomes. Switching from XML to nested JSON produces a 31% relative increase in attack success on Qwen3-14B memory-attack tasks, yet no wrapper family is consistently safest. (2) Weaker source separation is associated with higher memory-attack risk across renderings in three models spanning two model families. (3) Shifting external-message representations toward a trusted-source direction learned from neutral prompts without attack labels increases attack success by 9.2% on held-out Qwen3-14B memory-attack tasks, demonstrating a causal influence of source representations on attack susceptibility. These findings establish Protocol Rendering as a security-relevant design choice and support source separation as a contributor to rendering-dependent attack risk. https://anonymous.4open.science/r/ProtocolRenderingSafety-2848Code is available online.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.