acceptodds
Under review as a conference paper at ICLR 2027

Unwilling, or Unable? Identifying the Safety Cost of Tool Interfaces in LLM Agents

Abstract

Agent-safety benchmarks typically use refusal rate, but a model that declines because a request is harmful and one that declines because it has no way to act look identical in text. We show that this confound can materially change conclusions for several evaluated models. We use AgentHarm’s built-in benign twins which are matched, harmless tasks that share each harmful item’s structure and grader and construct a capability-adjusted estimator: a difference-in-differences that subtracts the twin’s refusal change, plus an item-level rate conditioning on a demonstrated capability match. The core estimators, audit framework, exclusion rules, and minimum-sample requirements were documented before the later hosted and local evaluation tiers; subsequent refinements are recorded in a timestamped deviation log, and reported results are generated from the released analysis pipeline rather than transcribed. We evaluate 16 hosted models and 8 locally served configurations spanning 10 organisations overall. On the 14 hosted models that support both interfaces, the mean measured effect moves from -0.319 to -0.188, 3 point estimates reverse sign, and 6 lose significance after correction. When both sides expose an executable action space, refusal-to-action conversion ranges from 3.5% to 77.8% across models. In separate multi-turn runs, 34.1% of GPT-4o’s demonstrated refusals become actions (20.0% fully completed), 58.9% of Llama-3.3-70B’s (21.1%), 62.5% of Mistral-Large-3’s (21.4%), 38.7% of Qwen3-32B’s (16.1%), and 6.8% of DeepSeek-V3.2’s (1.7%). On the three tested OpenAI models, the capability-adjusted graded-harm change is negative: raw harmful completion rises, but by less than completion on the matched benign controls. On the tested openweight configurations, the largest behavioural shift occurs when an executable action space first appears, while hosted native-channel effects are model-dependent. Harm remains linearly decodable at the decision token within framing on the Qwen-family states we instrument. A linear probe guard substantially reduces completed harm but exhibits calibration drift, and a dedicated decline task action restores refusal under a matched placebo on Qwen3-8B and hosted GPT-4o, while worsening harm on Gemma-4-E4B.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.