acceptodds
Under review as a conference paper at ICLR 2027

Surface Cues Are Not Enough: Towards Content-Grounded Defenses for Tool-Using Agents

Abstract

LLM agents combine reasoning with external tools to solve complex real-world tasks, but the same tools can also enable harmful actions, motivating guardrails to monitor agent behavior. However, existing evaluations construct harmful and benign tasks independently, so they often differ not only in safety-critical content but also in instructions and tools. This makes it difficult to tell whether guardrails truly identify the content that makes a task harmful or simply rely on these surface cues. To address this limitation, we introduce CoPair-Bench, a content-contrastive paired benchmark that systematically constructs harmful–benign task pairs while holding instructions and tools fixed, isolating the safety-critical semantic differences that determine harmfulness. Evaluations reveal that existing guardrails struggle to distinguish harmful tasks from their benign counterparts: they either allow both, providing negligible protection, or reject both indiscriminately, resulting in severe over-refusal. This behavior suggests that their decisions are driven largely by superficial cues, such as task instructions and tool names, rather than by the harmful objective itself. To mitigate this issue, we further propose Harm-Perception Shield (HP-Shield), a training-free self-exploration defense that prompts agents to first inspect content for safety risks and distills reusable safety experiences from iterative sandbox simulations. HP-Shield improves Safe-Utility by 14.48% on average compared with existing guardrails, enabling more accurate safety decisions without model retraining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.