TRACE-GUARD: Stateful Runtime Enforcement for Multi-Step Tool-Using Agents
Abstract
Multi-step safety risks in tool-using LLM agents are history-dependent: the same action can be safe after one sequence of actions but harmful after another. Existing defenses check actions using history or state. The challenge is to identify when earlier actions make the next action unsafe, before it is executed. We present TRACE-GUARD, a runtime defense that maintains a compact, structured safety state encoding safety-relevant information from earlier guard-approved actions. Before dispatch, it evaluates how a proposed action would change this state and decides whether to allow its execution. To evaluate runtime safety beyond end-task outcomes, we introduce RMSBench (Runtime Multi-Step Safety Benchmark), a 300-case executable benchmark for state-dependent runtime safety, with task-specific environment criteria and transition-level runtime evidence. Experiments on RMSBench across four models show that TRACE-GUARD substantially reduces task-criterion ASR (on DeepSeek-V3.2, mean ASR drops from 97.7% to 20.5%). Across 30 executable controlled-history pairs with identical final actions but different preceding histories, TRACE-GUARD distinguishes safe from dangerous cases before dispatch, demonstrating history-conditioned runtime enforcement beyond action-local checking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.