acceptodds
Under review as a conference paper at ICLR 2027

Learning Authorization for Tool Use

Abstract

An agent that learns to reject a file-sharing instruction hidden in a document must still share the file when its user asks. Training on the incident alone can suppress both actions, because the tool call is identical and its authorization is the relevant difference. We study this repair problem with matched preferences that keep the operation fixed while changing who requests it and whether the user has confirmed it. Preference optimization trains the agent to make the authorized call and answer or seek confirmation in the other contexts. A reference-KL penalty constrains changes on previously correct responses. On a controlled AgentDojo test, all three Qwen runs recover all 64 authorized actions while making no unauthorized state changes. Matched controls on natural-language requests reveal different effects by tool. On a fresh execution test, DPO selects calendar targets more accurately, while SFT produces more executable file-sharing calls. Llama recovers authorized execution with less reliable clarification. The retention studies find different effects across task sets, rather than a general benefit from KL. The results identify authorized-action recovery as a distinct target for agent repair: preventing the incident and completing the user's work must be learned and measured together.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.