Prompt Injection as an Information Channel: Detection Limits and Hierarchy Certificates
Abstract
Prompt-injection evaluations often conflate three distinct questions: whether benign and attacked output laws are distinguishable, how much attacker-controlled information can flow through generation, and whether a finite audit can certify a hierarchy-violation bound. We give a unified information-theoretic framework that separates these questions and connects each to an auditable quantity. We define benign and attacked model-induced output laws under an explicit observation model and characterize their optimal one-shot testing risk. We then replace an independent-token analysis with an autoregressive likelihood-ratio theorem: under auditable drift and martingale-tail conditions, detection error decreases with accumulated conditional information. We prove that attention entropy alone cannot provide an architecture-independent injection-capacity bound, and give an influence-aware alternative. Finally, protected-event data processing yields a hierarchy-violation certificate whose finite-sample form uses exact binomial coverage under a declared IID audit-unit model rather than zero-rate clipping. Across 18,432 Kaggle trials, attack-conditioned exact leakage spans 0.137–0.590 across four open-model configurations, while the best pair-conditioned oracle test reaches held-out AUC 1.000 and output-surface baselines remain substantially weaker. A controlled hierarchy LoRA reduces direct-override leakage but does not yield a confidence-certified increase in required KL divergence, separating empirical robustness from certification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.