When Agents Don't Take "No”: Curbing Guardrail Evasion while Keeping Utility
Abstract
Blocking an unsafe tool call does not necessarily block its harmful effect, as an agent can retry or reroute until subsequent actions produce the same outcome, even without an external attacker or explicit instructions to bypass the guardrail. We call this behavior guardrail evasion. Evaluating four existing guardrails on two agent-security benchmarks, RedCode and SABER, we find that every guardrail either allows risky actions to proceed or hinders legitimate task execution. These failures arise primarily because agents either disregard or misinterpret the guardrail's verdict. We develop \sysname, which derives effect-oriented rules from reported incidents, hardens them against effect-preserving rewrites through counterexample-guided refinement, and augments each denial message with an agent-steer prompt that curbs evasion without disrupting legitimate work. With one rule set that transfers across agents, benchmarks, and harnesses, \sysname lies on the Pareto front of safety, utility, and cost against four baselines in all five settings we test. On RedCode, for example, \sysname cuts the attack success rate of a DeepSeek-V4.1-Flash agent in OpenCode from 49.7% to 7.8%, below every baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.