From Traces to Explanations: A Structured Framework for Auditing AI Agent Behavior
Abstract
AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did, how its behavior unfolded, and what evidence supported it. However, traditional Explainable AI (XAI) methods fall short of providing the process-level transparency required for such interactive, multi-step systems. To address this gap, we introduce a trace-grounded structured representation of agent execution that reconstructs behavioral trajectories from lengthy execution traces and provides an auditable basis for faithful natural-language explanations. The representation explicitly separates behavioral commitments from the internal and external evidence, enabling the identification of behavior-evidence mismatches. Relying solely on execution traces, our framework generalizes across agent architectures and environments. Human and automated evaluations across multiple benchmarks and architectures show that our framework produces high-quality, trace-faithful explanations while reliably identifying unsupported claims, unjustified actions, and evidence gaps, outperforming direct LLM-generated explanations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.