acceptodds
Under review as a conference paper at ICLR 2027

Tracing Paths to Risk: Automated Safety Evaluation for Tool-used Agents via Risk-Path Exploration

Abstract

LLM agents increasingly act through external tools, extending safety risks beyond generated content to actions in the environment. Existing agent safety evaluations often rely on manually constructed or fixed test cases, which sparsely sample the execution space of each risk. However, the same risk may arise from different tool-use trajectories, environment states, and trigger conditions, leaving many feasible failure modes unexplored. To address this limitation, we propose a risk-path-driven framework for automated agent safety evaluation that treats each risk category as a family of reachable execution paths. Given an agent and a target risk, the framework constructs a compact risk-relevant graph by pruning irrelevant tools, enabling efficient search for feasible and diverse risk paths that are instantiated into executable test cases. It further refines failed cases using execution feedback, allowing test cases to adapt to the target agent’s behavior. During evaluation, a verifier determines whether the target risk has been triggered using tool-call traces and environment-state evidence, rather than relying solely on the agent's self-reports or final responses. Experiments across 349 task scenarios, 7 agent configurations, and 8 risk categories show that our framework achieves a 91.25% average risk discovery rate and uncovers diverse path-dependent safety failures that fixed test cases may miss.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.