acceptodds
Under review as a conference paper at ICLR 2027

To Call or Not to Call: Mastering Tool Admission in Language Agents

Abstract

Large Agents increasingly call external tools, yet most safety research examines only what follows a tool call. The critical prior decision of whether to call the tool is unexamined and can heavily influence downstream safety. Across multiple open-weight models, the unweighted model-average harm rate after admitted harmful requests reaches a strikingly high level. To identify what controls this decision, we apply a factorial intervention to the input. Within our generated construction, counteracting the requested-operation component produces the largest admission contrast. A linear direction estimated from this contrast influences held-out admission in most tested models; effects on operation-neutralised requests in some models show it is not an operation-specific mechanism. We then test whether written refusal is behaviorally and representationally coupled to the same decision. Surprisingly, written refusal accompanies a substantial portion of admitted harmful calls. Natural language refusal training reduces unsafe admission in models that retain benign tool use, showing that response-level safety training can effectively transfer to the call decision. These results establish a partial, model-dependent relation between written refusal and tool admission. They motivate direct evaluation and supervision of tool admission, proving we can no longer treat text refusal as a simple proxy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.