acceptodds
Under review as a conference paper at ICLR 2027

SafetyBench: Evaluating Agent Safety Under Benign Asks

Abstract

Modern agents are highly capable. They can execute complex workflows involving communication, documents, and operational tools. Oftentimes, these tasks require the agent to interpret boundaries that the user request leaves implicit: what information to share or what actions are authorized. For example, an agent asked to schedule off-boarding discussions should arrange the meeting without prematurely revealing an employee's resignation through a shared calendar invitation. To evaluate how reliably agents make these judgments, we build SafetyBench: a collection of 200 simple operational tasks across 8 workspace environments designed to evaluate safety under delegation. Each task is a legitimate, benign user ask, but provides opportunities for unsafe actions. We evaluate frontier models across multiple agent harnesses and distinguish between safe completion, deference to the user, and unsafe action.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.