acceptodds
Under review as a conference paper at ICLR 2027

Control-Plane Prompt Injection: Forging Authority in LLM Content Moderation

Abstract

Large Language Models (LLMs) deployed in content moderation increasingly process multi-channel inputs where untrusted user data can disguise itself as system infrastructure. We formally define and systematically evaluates this attack as control-plane prompt injection (CPPI), a vulnerability where attackers forge audit caches, label assignments, or formatting constraints within body text or user comments to hijack the control authority of the target model without altering the system prompt. CPPI exploits ambiguity between data and control signals, and causes models to copy forged labels or execution directives embedded in ordinary content. Extensive experiments show that CPPI successfully disrupts six representative defense measures on five mainstream LLMs. Our study identifies control-content confusion as a fundamental vulnerability of modern LLM moderation systems. We suggest that future prompt injection research should move beyond instruction conflicts toward understanding how models assign authority to heterogeneous input sources.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.