Risk Exposure Does Not Mean Unsafe Action: Evidence-Dynamics Alignment for Agent Safety
Abstract
Large language model (LLM)-based agents increasingly operate through multi-step interactions with external environments and tools, creating new challenges for trajectory-level safety monitoring. A key difficulty is that risk exposure does not necessarily imply unsafe action: an agent may encounter risky contexts while ultimately behaving safely, yet existing LLM-based monitors can conflate the two and produce false alarms. We find that this failure does not simply arise from missing safety information. Frozen monitor representations retain distinguishable evidence of risk exposure and unsafe enactment, but such evidence is not fully reflected in native decisions. Based on this observation, we propose Evidence-Dynamics Alignment (EDAlign), a lightweight post-hoc framework that tracks Exposure and Enactment evidence across model depth and aligns final safety decisions with their evolving evidence state. Experiments across agent safety benchmarks and LLM backbones show that EDAlign reduces exposure-induced false alarms while maintaining strong detection of unsafe behavior. Comparisons with direct unsafe probes and layer-range ablations further support the value of structured Exposure–Enactment evidence and its evolution across model depth.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.