Control-Plane Prompt Injection: Forging Authority in LLM Content Moderation
Abstract
Large Language Models (LLMs) deployed in content moderation increasingly process multi-channel inputs where untrusted user data can disguise itself as system infrastructure. We formally define and systematically evaluates this attack as control-plane prompt injection (CPPI), a vulnerability where attackers forge audit caches, label assignments, or formatting constraints within body text or user comments to hijack the control authority of the target model without altering the system prompt. CPPI exploits ambiguity between data and control signals, and causes models to copy forged labels or execution directives embedded in ordinary content. Extensive experiments show that CPPI successfully disrupts six representative defense measures on five mainstream LLMs. Our study identifies control-content confusion as a fundamental vulnerability of modern LLM moderation systems. We suggest that future prompt injection research should move beyond instruction conflicts toward understanding how models assign authority to heterogeneous input sources.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.