acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Refusal Failures Under Tool Schemas

Abstract

LLMs are aligned to refuse harmful requests, but they often comply with the same requests once given a tool schema. This gap is well documented, and training on tool-use data can reduce it, but what happens inside the model remains largely unexplored. We show that the failure lies not in recognizing harm but in acting on it. To read the tool decision, we force the model to start a tool call and compare how likely it is to name a decline tool against the offered execution tools, a quantity we call the decision margin. Across four instruction-tuned models from three families, a harm direction stays linearly decodable under the schema, and steering the last request token along it raises the decision margin beyond random directions, yet the action changes on only some models. We then ask what safety training changes, and fine-tune each model on 1,000 requests to refuse, with and without the tool schema. Training inside the schema lowers harmful tool calls, while training without it largely fails or stops the model from calling tools at all. A smaller test suggests the same split between recognizing harm and acting on it under JSON, Shakespeare and list output styles. Safety evaluation and training for tool-using agents therefore need to happen inside the tool format. More broadly, our findings suggest that safety alignment is bound to the format it was trained in, which calls for alignment methods whose safety transfers across formats.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.