acceptodds
Under review as a conference paper at ICLR 2027

Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment

Abstract

Large Language Models (LLMs) are increasingly deployed as agents that operate in real-world environments, introducing safety and security risks beyond linguistic harm through tool use, external transactions, and multi-step decision making. Existing agent safety evaluations construct risk-oriented tasks that lack a hierarchically organized taxonomy of attack surfaces, rely on predefined harm categories rather than deployment-grounded safety rubrics, and evaluate agent behavior in isolation rather than during realistic, long-horizon task execution. To address these limitations, we propose Risky-Bench, a deployment-grounded safety evaluation layer for life-assist agents. Built on top of existing life-assist task environments, Risky-Bench derives context-aware safety rubrics from domain-agnostic safety principles and systematically probes whether agents violate these rubrics while pursuing otherwise benign user goals, under adversarial perturbations across hierarchically organized attack surfaces with explicitly separated threat assumptions. Applied to simulated life-assist scenarios covering food delivery, in-store assistance, and online travel booking, Risky-Bench evaluates seven representative contemporary agents across 750 tasks, revealing average attack success rates ranging from 23% to 58% and exposing systematic failure patterns including threat-model-specific vulnerabilities and model-dependent effects of explicit reasoning. Comparisons with AgentHarm and Agent Security Bench further show that Risky-Bench provides complementary, non-redundant safety signals. Our code and data are available at https://anonymous.4open.science/r/Risky-Bench-D892.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.