acceptodds
Under review as a conference paper at ICLR 2027

Graph Verification for LLM Agent Safety

Abstract

As large language models (LLMs) evolve into more capable agentic systems, their safety risks become increasingly diverse and consequential, calling for a reliable and verifiable safety assessment. Existing verification methods typically encode safety requirements into explicit rules or specifications, but the evidence needed to evaluate them remains scattered across complex execution trajectories. In this work, we first analyze common agent risks and characterize agent safety as a set of conditions on the sources and effects of each agent action. To better support the assessment of these conditions, we introduce **VEGAS**, a **VE**rifiable **G**raph for **A**gent **S**afety that organizes the complex trajectory contents into a unified graph representation, and formalizes the safety conditions as explicit graph properties. Furthermore, we also develop GraphGuard, a runtime realization of the above safety formalization that focuses on optimizing practical efficiency and enabling benign-task continuation after safety blocking. Experimental results show that our method substantially improves agent safety by reducing ASR of DeepSeek-V4-Pro from 45.3% to 0.0% on AgentDojo, while still maintaining a high benign-task performance. Our work demonstrates the potential of using graph representation to provide a structured and verifiable abstraction for agent safety.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.