acceptodds
Under review as a conference paper at ICLR 2027

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails

Abstract

LLM-based guardrails have emerged as an effective defense against prompt injection and jailbreak attacks in autonomous agents. However, we reveal that the very reasoning and task-following capabilities enabling this protection introduce a novel vulnerability: attackers can inject crafted data to trap the guardrail in extended reasoning loops, effectuating a systematic denial-of-service (DoS) attack. To systematically expose this threat, we design a beam-search optimization framework that crafts natural-language payloads to maximize guardrail reasoning length, utilizing an LLM proposer guided by a strategy bank. Based on the observation of guardrail's schema-following nature, we also provide another attack framework driven by mechanism-aware structural mutations with less computational load. The attack efficacy is evaluated in two parts. First, standalone evaluations cover 10 target guardrail models across safety templates and agent benchmarks. Payloads optimized on an open-source surrogate successfully transfer to leading model backbones including Claude, GPT and Gemini, achieving a 13–63 length amplification. Second, evaluations in web, code, computer-use and multi-agent deployments show up to 148 guardrail-latency amplification. In a shared-guardrail supervisor experiment with sequential access, attacked workers may reduce throughput by 23.3% and delayed benign workers by 113–146 seconds. These findings motivate explicit reasoning-cost bounds and reliable verdict recovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.