AgentXRay: Grounding Agent Safety in System Evidence
Abstract
Computer-use agents translate natural-language instructions into consequential operating-system actions, including file access, command execution, permission changes, and network communication. Existing safety evaluations primarily rely on model responses, UI trajectories, screenshots, or tool calls. Although these traces reveal task context and apparent intent, they do not reliably establish whether attempted actions succeeded or what effects occurred in the underlying system. We introduce AgentXRay, a joint-evidence framework that pairs agent trajectories with independently collected Linux audit log. To our knowledge, AgentXRay is the first framework to pair synchronized agent trajectories with kernel-observed OS provenance and human annotations. Our evaluation uses 53 controlled executions spanning deliberate user misuse, prompt injection, model misbehavior, and benign activity with human annotation. We normalize each paired trace into an Agent-System Evidence Graph (ASEG), which represents agent actions, audit events, processes, affected objects, and temporal-semantic associations across the two evidence layers. An evidence-grounded LLM judge uses these representations to assess safety and task completion. Our evaluation shows that this alignment strategy is effective for agent safety detection, with particularly clear gains on prompt-injection executions. More broadly, the results suggest that explicitly structuring and relating agent actions to observed system effects is more effective than simply concatenating the two evidence sources.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.