When and Why Do Agents Become Unsafe? A Trajectory-Aware Benchmark for Multi-Turn Agent Safety Evaluation
Abstract
By leveraging external tools, large language model (LLM) agents have achieved substantial capability gains, but this tool-augmented interaction also introduces more safety risks that can emerge at any stage of a multi-turn interaction. However, existing safety benchmarks mainly focus on whether an agent's final behavior is unsafe, without considering when and why agents become unsafe. To address this limitation, we introduce \method, a benchmark for fine-grained safety evaluation of LLM agents under multi-turn attacks. Specifically, \method comprises 1,100 multi-turn attack cases ( in English, in Chinese) across scenarios, covering both content and behavioral risks. These cases are constructed through dynamic agent interaction using multi-turn attack strategies with tool categories, accompanied by detailed human annotation. We then perform multi-turn rollouts of target agents on the constructed cases for evaluation. Furthermore, we propose a novel evaluation framework for risk tracing, including progressive safety diagnoses: Safety Judgment determines whether an agent generates unsafe content or takes unsafe actions, Risk Localization identifies the first risk step, and Capability Diagnosis characterizes the failure along four safety capability dimensions using evidence from the localized step and its context. Overall, \method connects a trajectory-level safety judgment to the specific behavior and capability failure it reflects. Experiments on LLMs across agent harnesses reveal that Claude Sonnet 4.6 shows the strongest safety, while Llama 3.3-70B shows the weakest. Further analysis shows that task scenario disguise is the most effective attack strategy, and agent safety failures are mainly attributable to risk understanding and tool-use boundary control. The project page is https://mtrisktrace.github.io/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.