LLM Agent Refusal Can be Harmful: An Availability Attack through MITM Tampering on Remote MCP Services
Abstract
Safety-aligned LLM agents refuse sensitive contexts more readily, including the return values of tool calls. We show that this refusal behavior is itself an attack surface: fabricated security alerts embedded in MCP return values induce agents to brick legitimate user requests, constituting an availability attack on agentic harnesses. We propose the ARB attack, which appends covert fake alerts to MCP communications that the attacker hijacks or relays as a man in the middle; the attack targets requests that depend on a remote MCP service and cannot be bypassed by local tools. We formalize the attack with a refusal-probability decomposition into tool-call probability, warning sensitivity, and post-alert verification capability, and validate it on 10 mainstream models across 3 harnesses: ARB exceeds 50% ASR in 90% of the 30 model–harness configurations. We further construct NBMR, the first MCP dataset designed around non-bypassability (8 categories, 160 docker-deployable scenarios), and propose AEPG, an MCTS-based defense that synthesizes AGENTS.md rule files; the synthesized prompt raises executed turns on the attacked agent from 18.8% to 75.0% and scene completion from 12.5% to 75.0%. Stress tests show that the induced refusal persists through the full repetition budget for the most aligned model families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.