WardenGuard and Warden-Bench: A Multichannel Defense and Benchmark for Indirect and Agentic Prompt Injection
Abstract
Prompt injection is the leading security risk for language-model agents, yet existing benchmarks cover narrow attack classes and existing detectors either flatten the agent's context into a single string or fine-tune on one attack distribution, causing high false positives or poor generalization. The proposed work makes two contributions. First, Warden-Bench, a 23,601-row benchmark consolidating twelve attack tiers (T0-T11), direct overrides, indirect injection via tool outputs and retrieved context, multi-turn fragmentation, and encoding evasion, plus a 5,899-row benign split for false-positive measurement. Second, WardenGuard, a defense exposing six trust-annotated context channels to a two-stage judge cascade: a local open-weight model settles clear cases, deferring to a frontier judge when an input is unfamiliar or escalation is needed, behind an attestation gate and a write-time memory-integrity monitor, preceded by a deterministic 2 ms decoder covering thirteen of fourteen encoding schemes, with every decision hash-chained for audit. The proposed work evaluates across Warden-Bench and four external benchmarks (41,949 cases total), WardenGuard attains mean F1 0.911 and precision 0.981 externally, and is the only defense meeting F10.90, FPR5%, and FNR15% on Warden-Bench. Across a chained residual attack matrix (442,550 combinations), it reaches the lowest mean residual ASR (1.95%), significantly below every competitor on all twenty cells at 11.4 fewer false positives, a margin that persists under an open-weight judge backbone. Against a query-budgeted adaptive attacker, it admits 34.0% of a stratified held-out sample versus 68.7% and 70.0% for the strongest baselines, a separation that survives excluding every case the attestation gate refuses, indicating that trust-annotated context-channel separation is a practical, backbone-agnostic path toward deployable prompt-injection defenses. The code is available at: https://anonymous.4open.science/r/WardenGuard-07F4/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.