CROSS: Coupled Recurrence of Latent-Symbolic Security-State for Runtime Safety Monitoring of AI Agents
Abstract
Long-horizon tool-using agents can encounter safety failures whose decisive evidence is distributed across individually innocuous actions. Earlier observations may change the safety of a later tool call, while successive actions may diverge from the user's authorized objective. Monitoring such failures requires trajectory memory, but repeatedly encoding the full history increases computation, and a fixed local window discards earlier evidence. We propose **CROSS** (**C**oupled **R**ecurrence **O**f latent-symbolic **S**ecurity **S**tates), an action-level monitor that couples recurrent neural state with a symbolic finite-state machine. The neural state summarizes contextual cross-step evidence; the symbolic state retains explicit security facts and determines auditable interventions. Symbolic state conditions the neural update, and neural predicates drive the symbolic transition, while the neural observation window remains bounded. We evaluate CROSS on synthetic risk-propagation and goal-shift tasks and on StepShield, AgentDojo, and Upward Deception. With a 1.7B backbone, CROSS achieves the highest F1 among all evaluated guards, including a baseline with a 7B backbone, on the synthetic dataset and all three public benchmarks. These results demonstrate the effectiveness of coupled recurrence for cross-step safety detection in the evaluated settings. Ablation studies further support the value of complementary security states and contextual evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.