Hijacking the Instruction Hierarchy in LLM Agents: From Semantic Injection to Persistent Compromise
Abstract
Major LLM providers have introduced instruction hierarchy training to strengthen model security overall against prompt injection. However, with the rise of LLM agents and skill-sharing ecosystems, third-party skill text has become a new attack surface: its benign metadata enters high-authority positions through the agent's own placement after installation, and the model's hierarchical instruction following transfers to the skill body along the metadata-body mapping, which attackers can exploit in agent settings to compete against the model's own safety alignment. We also find that mainstream LLMs readily detect explicitly malicious code but poorly recognize natural-language text carrying equivalent malicious intent. These two factors lead the agent to generate and execute malicious scripts aligned with attacker intent, and even to write malicious textual instructions into local memory files whose effects persist after the skill is removed. This elevates indirect prompt injection from transient behavior disruption to substantive and persistent system compromise. We propose HierJack and present a comprehensive study of how LLM agents expose the instruction hierarchy to hijacking. Systematic evaluation across 16 LLM-agent combinations yields an average one-shot attack success rate (ASR) of 88.46% and an average persistent ASR of 69.57%. Porting the payloads of InjecAgent and ChatInject to skill-based delivery vehicles raises their ASRs by 46.4 and 36.9 percentage points over the original external-data delivery. These results show that the LLM agent shifts the security boundary from model-side to agent-side privilege granting. Once an agent grants third-party content high authority, model compliance becomes an attack amplifier when it outcompetes the model's safety alignment on specific content. Safety alignment must therefore extend to both the optimization of privilege tiering and the recognition of semantically encoded malicious intent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.