Recovering Structural and Semantic Constraints for LLM Agent Anomaly Detection
Abstract
Large language model (LLM) agents increasingly rely on external tools to accomplish complex tasks, exposing their execution processes to failures and attacks that may alter tool invocation sequences or their semantic dependencies. Existing defenses often focus on predefined rules or individual interactions, making it difficult to identify diverse and previously unseen anomalous behaviors. In this work, we demonstrate that normal agent executions themselves can provide reusable constraints for anomaly detection. We present TraceAegis, a framework that recovers structural and semantic constraints from historical agent execution traces. TraceAegis first organizes recurring tool interactions into hierarchical execution units and recovers structural constraints that characterize valid invocation flows. It further summarizes semantic constraints over these units to capture the contextual conditions under which tool interactions remain consistent with the intended task. At detection time, TraceAegis detects anomalous executions by identifying violations of either structural consistency or semantic coherence, allowing it to capture both unseen execution paths and previously observed paths with abnormal semantic dependencies. To evaluate the effectiveness of TraceAegis, we introduce TraceAegis-Bench, a benchmark containing 2,600 benign and 600 abnormal execution traces across healthcare and corporate procurement scenarios, and further evaluate TraceAegis on execution traces collected from an internal red-team exercise at a technology company. Across five LLM backbones, TraceAegis achieves higher F1 scores than the corresponding LLM-based baseline detectors in both scenarios. With DeepSeek-V3, it achieves F1 scores of 0.965 for clinical triage and 0.932 for corporate procurement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.