Tracing Paths to Risk: Automated Safety Evaluation for Tool-used Agents via Risk-Path Exploration
Abstract
LLM agents increasingly act through external tools, extending safety risks beyond generated content to actions in the environment. Existing agent safety evaluations often rely on manually constructed or fixed test cases, which sparsely sample the execution space of each risk. However, the same risk may arise from different tool-use trajectories, environment states, and trigger conditions, leaving many feasible failure modes unexplored. To address this limitation, we propose a risk-path-driven framework for automated agent safety evaluation that treats each risk category as a family of reachable execution paths. Given an agent and a target risk, the framework constructs a compact risk-relevant graph by pruning irrelevant tools, enabling efficient search for feasible and diverse risk paths that are instantiated into executable test cases. It further refines failed cases using execution feedback, allowing test cases to adapt to the target agent’s behavior. During evaluation, a verifier determines whether the target risk has been triggered using tool-call traces and environment-state evidence, rather than relying solely on the agent's self-reports or final responses. Experiments across 349 task scenarios, 7 agent configurations, and 8 risk categories show that our framework achieves a 91.25% average risk discovery rate and uncovers diverse path-dependent safety failures that fixed test cases may miss.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.