Guard-and-Guide Runtime Safety for Long-Horizon LLM Agents
Abstract
Large language model (LLM)-driven agents are increasingly deployed in long-horizon tasks, where the harms of individually low-risk actions can accumulate over time and eventually lead to safety failures. Existing parametric alignment methods and non-parametric memory mechanisms struggle to detect subtle risks and provide timely protection. To address this challenge, we propose G2A (Guard-and-Guide), a framework for safe long-horizon agent execution. A lightweight Guard scores candidate actions in the current interaction context. When a risk score exceeds a specified threshold, a procedural Guide selects and executes a reusable Safety Skill. To detect context-dependent risks, G2A introduces Trajectory-Contrastive Safety Distillation (TCSD), which contrasts the same candidate action under matched safe and unsafe trajectory contexts and distills the resulting evidence into the Guard. The Guide organizes Safety Skills through explicit activation, intervention, and termination conditions, enabling state-aware skill activation, switching, termination, and fallback. It also records the scenario and action risk type at each use to support evolving skill reuse. Experiments on ATBench-C, AgentS4D, and HINTBench show that G2A improves long-horizon safety, enables fine-grained risk localization and cross-scenario skill reuse, and preserves agents’ task completion capability.Our code is available at https://anonymous.4open.science/r/G2A-2228.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.