From Security to Safety: Towards Non-Adversarial Risks of AI Agents
Abstract
As AI agents increasingly enter production workflows and take autonomous actions, ensuring their safe and reliable operation becomes critical for real-world deployment. While existing research has extensively studied adversarial threats in *Agent Security*, *Agent Safety* risks arising in non-adversarial settings remain insufficiently clarified and explored. To bridge this gap, we clearly clarify *Agent Safety* risks as *non-adversarial risks*: harms caused by agents autonomously taking unsafe actions under benign user intent and without third-party attackers. Building on this clarification, we introduce *risk model* to capture the causal pathways from risk triggers through inherent agent failures to unsafe actions, and propose a unified framework for organizing these models. We further develop a three-stage risk task generation pipeline that translates abstract risk models into high-quality executable tasks, and use it to construct the scalable *NARA-Bench*, which currently comprises 1,800 executable risky tasks spanning 6 risk models, 12 scenario scopes, 45 tools, and 7 risk-outcome categories, enabling end-to-end, high-fidelity evaluation of *Agent Safety*. Systematic evaluations show that current agents remain susceptible to severe and diverse risky behaviors even in benign environments. Our findings reveal substantial non-adversarial safety challenges and provide a foundation for evaluating and improving the reliability of agents in real-world deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.