SafetyBench: Evaluating Agent Safety Under Benign Asks
Abstract
Modern agents are highly capable. They can execute complex workflows involving communication, documents, and operational tools. Oftentimes, these tasks require the agent to interpret boundaries that the user request leaves implicit: what information to share or what actions are authorized. For example, an agent asked to schedule off-boarding discussions should arrange the meeting without prematurely revealing an employee's resignation through a shared calendar invitation. To evaluate how reliably agents make these judgments, we build SafetyBench: a collection of 200 simple operational tasks across 8 workspace environments designed to evaluate safety under delegation. Each task is a legitimate, benign user ask, but provides opportunities for unsafe actions. We evaluate frontier models across multiple agent harnesses and distinguish between safe completion, deference to the user, and unsafe action.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.