acceptodds
Under review as a conference paper at ICLR 2027

Why Do Defenses Fail? Counterfactual Diagnosis of LLM Agent Safety

Abstract

A lower attack success rate (ASR) in LLM agents can reflect selective protection or impaired legitimate behavior. We introduce Three-Axis Safety Diagnosis through Counterfactual Measurement and Evaluation, or TASD-CME, a white-box protocol that tests a proposed explanation by pairing a prespecified diagnostic test with a targeted internal intervention at the same frozen decision point. The protocol asks whether external directives receive undue influence, harmful use can be suppressed while preserving benign capability, and safe actions are preferred when harm is recognizable. Local support requires a qualifying failure signature, verified intervention delivery, the intended target response, behavioral improvement, retained legitimate performance, and effects beyond paired controls. Across five attack families and two models, prespecified interventions lower first-action ASR point estimates in all ten settings, but only four satisfy this joint rule. For tool-chain manipulation, Gemma and Qwen reduce ASR by 58.00 and 50.53 percentage points, respectively. Qwen meets the joint requirements; Gemma's diagnostic comparison fails the prespecified balance check despite its larger behavioral gain. In Qwen, the intervention retains 98.0% of baseline legitimate-task utility and exceeds the random control's ASR reduction by 44.13 points. TASD-CME separates behavioral gains from support for their proposed local explanations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.